Managing releases
A new version of PHP, of the agent, or of the panel itself does not go to every server at once. It walks: a soak on the beta channel, then a canary machine, then the fleet in waves, with a wait between each pair and a health check that can
Written for: Administrator
A new version of PHP, of the agent, or of the panel itself does not go to every server at once. It walks: a soak on the beta channel, then a canary machine, then the fleet in waves, with a wait between each pair and a health check that can stop the whole thing. This page is how you drive that from the panel, and how to read what it tells you.
Only administrators see this. Resellers and end customers never see which versions exist, let alone which machine is on which one.
Open Servers → Releases.
The three questions the board answers
The page is laid out in the order you actually ask them.
1. May a wave start right now? The banner at the top. It is green when the door is open and amber when it is not, and when it is not it says why in a full sentence — a Saturday, a public holiday, or an incident somebody opened.
2. What is waiting on me? The pipeline. Every version with a badge for its risk class, its channel, and how far its soak has run.
3. Which machines would it touch first? The wave groups, at the bottom. Drag a server from one card to another to change the order the fleet is updated in.
Risk classes
The badge on every release decides everything about how it travels.
- minor-runtime — a PHP patch, a LiteSpeed build. Nothing on the node keeps state about it, so a downgrade is a real answer, and the panel promotes it by itself once its soak comes up clean.
- platform — the agent, or the panel. It promotes only when a person clicks, its canary is at least three machines including a busy one, and the panel rolls forward rather than back: a database that has migrated does not un-migrate because a binary went backwards.
- security-critical — a fix that cannot wait. The timers compress. Nothing else does.
An unknown package is treated as platform, which is the conservative answer. You can override the class of a package under Advanced → risk classes if the name-based rule gets it wrong.
Approving a release
A release whose class needs a person appears with the state awaiting_approval, and the banner at the top of the page counts them. Open it to see the evidence:
| Soak | how many days it has been on beta, against how many its class asks for |
| Canary | how many canary machines took it, and how many passed their gate |
| Error logs | the rate customers' error logs grew, as a multiple of before |
| Exposure | how many synthetic requests really succeeded on it |
These are the same four numbers the panel promotes a minor-runtime release on. There is no second, hidden set.
If the soak is clean, Approve moves it to the next channel. If it is not, the button reads Approve anyway and requires a reason — which goes into the promotion ledger with your name on it, and stays there.
Reject also needs a reason. It is the only record of why a release did not go out.
What is in this release
Above the evidence there is now a link: What is in ‹version›. It takes you to the release notes for exactly that release of the node agent — the software this rollout is about to put on your machines.
On this screen you are about to send a build to customers' servers. Until recently the version number was the whole answer to "what changed"; now it is there in plain words, grouped by category, with the details one click away.
The link appears only for corecp-agent and only when notes for that version exist. The other packages in the pipeline are other people's software and we publish no notes for them; an in-between build (a number with ~dev in it) does not have them yet.
You can open the list without a release too: Changes in the sidebar, and switch between Panel and Node agent at the top.
Wave groups
test is on beta and holds the operator machines with synthetic customer sites on them. wave-1, wave-2 and wave-3 are on stable, in blast-radius order.
Two things the panel insists on:
- Wave 1 needs a genuinely busy server. The card shows how many hosting accounts each machine carries and marks the busy ones. A first wave of idle machines exercises none of the code paths the release will actually meet, so the board warns when wave 1 has none — and a rollout planned that way is refused outright.
- A new server lands in a late wave by itself, round-robin, as soon as you open this page. Drag it forward when you know what it carries.
A group can hold itself back with a channel pin or a version ceiling — a customer who wants to stay a release behind. Every class honours that except security-critical, which reaches them anyway unless you opt the group out explicitly, and that opt-out is recorded with your name and your reason.
What stops a rollout
While a wave runs, each machine is checked after its own update — not "did apt exit zero", but whether its customers are as well served as they were ten minutes earlier: the websites still answer, their error logs are not filling up, mail is not queueing.
When that check goes red the panel does four things and then stops:
- rolls that machine back to the version it had,
- halts the rollout, leaving the machines it never reached marked as skipped,
- opens an incident, which blocks every further wave until you close it,
- sends the paging mail.
It never retries. A rollout that retried a failed machine would flap, and every cycle of that is another outage for the sites on it. The fleet stops where it stands and waits for you.
The panel always goes first
The panel may never be older than the servers it manages. So every rollout starts with the panel itself, and only then do the servers follow, wave by wave.
The rollout plan
Releases → Rollout plan holds the order. The first step is always The panel: that row has no buttons, because it cannot be moved or removed. Below it are the waves you name:
- Click Load production template for the production layout: first the test server
s1, thenn1,b1,d1andm1together, andw1last. Or click Add wave and fill in a name and the servers yourself. - Type servers as names, separated by commas:
s1or the full name. The arrows move a wave up or down — the first wave cannot move up, because the panel is above it. - Click Save plan. The panel asks for your second factor.
After saving, each wave shows which servers its names match. A name that matches no server, or more than one, is shown in orange. Servers that are in no wave are listed under the plan: a rollout from the plan does not reach them.
A plan with a wave before the panel is refused — also when it arrives through the API.
Start rollout from plan starts a rollout over the waves in their order, with the first wave as the canary. The panel refuses while a server could install a newer CoreCP version than the panel itself runs: update the panel first.
Updating the panel itself
A panel installed as a package updates itself from its own package source. You do this on the panel server, as root:
- First a backup of the panel database is made.
- The new version applies the database changes that do not get in the way of the current version, while the current version keeps running.
- The package is swapped, and the panel must answer with the new version within three minutes, and keep running.
- Only then does the last database step run.
If step 2 fails, the old version is still running and there is nothing to put back. If step 3 fails, the panel puts the previous package back by itself. Either way you get a notification. As long as the last database step has not run, the previous package can go back.
In the rollout plan, The panel shows the version that is running and how the last update went.
A server that was already unhealthy
Before a server is updated, the panel asks it how it is doing. If it already reports itself unhealthy — because it is still waiting for a restart, for example — it is not updated and the rollout stops with the sentence:
this server was already unhealthy before the rollout: pending updates: reboot required
Nothing was installed on that server. First fix what the sentence names (restart the server, for example), then start a recovery rollout. If the panel could not read the server's health, the server is not updated for the same reason.
From a terminal
On the panel server, as root:
# Once: let apt use this panel's own package source
corecp-panel self-update source
# Update to 0.50.1: backup, swap, check, finish
corecp-panel self-update --to 0.50.1
# Or stop before the last database step, and finish or roll back later
corecp-panel self-update --to 0.50.1 --hold-contract
corecp-panel self-update --finish
corecp-panel self-update --rollback
# How the last update went
corecp-panel self-update statusUpdating one server now
Not every update is a rollout. When one machine is behind — you have just put it back, it was off for a day, or you want it ahead of the rest — there is Update now, on the server page and on its Components tab.
What that button does, and above all what it does not:
- it installs what that machine's own channel offers now. It does not move the channel; that is a separate decision;
- it does not remove a rollback pin. If one is in force it stays, and the screen says the machine is being held back deliberately;
- it does not wait for the maintenance window. That window is there for the unattended timer; you are pressing the button yourself.
Afterwards there is one sentence: upgraded, already up to date, held at a pin, or declined, because …. The same sentence is kept on the task, so three hours later it is still readable what that update did — including by a colleague who was not there. From a terminal this is the same operation:
ssh root@stck1.corecp.dev 'corectl update'Servers outside the version window
The panel and the agent on a server talk over a fixed contract. That contract grows — an operation is added, a field is added — and every growth is a new protocol revision. The panel supports its own revision and the two before it. That is exactly why a rollout may be halted halfway: a fleet where half the machines have not been updated is a normal state, not an incident.
When a server falls outside that window you see it in two places:
- on Servers, a bar above the list naming the machines it is about;
- on the server's own page, a bar saying what is wrong and what you do about it, plus a marker beside the name.
The four cases, and what you do:
| What it says | What it means | What you do |
|---|---|---|
| does not report a protocol version | the agent predates CoreCP 0.20.5 and cannot answer the question | update the agent on that server |
| too old for this panel | the agent is more than two protocol revisions behind | update the agent on that server |
| newer than the panel | the server is ahead of the panel, usually after the panel was rolled back | update the panel first |
| a different protocol | the panel and the server are on different major versions | bring them to the same major, panel first |
You update with Update now on the server page, with a rollout, or from a terminal:
ssh root@stck1.corecp.dev 'corectl update'Warn first, refuse later
By default the panel only warns: it carries your changes out and tells you which servers are behind. Whoever runs the platform can turn that into a refusal — first on one pilot server, then fleet-wide. That lives in the panel's configuration and not in a screen, because making a fleet half deaf is an act with a maintenance window around it.
When the panel does refuse, two things always keep working: reading the server, and the update itself. Otherwise a refused server would be a server you can no longer repair. Only changes — creating an account, adding a domain — are refused, with the reason and the recovery step beside them.
What changed between two agent versions is in the agent's release notes, and the refusal points at them itself.
A server that cannot report something yet
A server inside the window is allowed to be missing something: the contract grows by adding, so an agent two revisions back answers everything it knows and says nothing about the rest. That is not a fault and the panel does not treat it as one.
Where it shows, it is spelled out with the server and the version. On the certificate detail, for instance: click a certificate on a server that does not report the request history yet and the panel names that server, the protocol version it does speak, and that updating it makes the history visible. The list itself is unchanged — everything that server can report is still there.
That distinction is why it is worded that way: "this machine has not recorded an attempt" and "this machine cannot report it" are two different sentences, and only one of them names something to do.
What is on a server: Components
A server's Components tab is its inventory: per kind — agent, PHP (FPM), PHP (LiteSpeed), database — which version is installed, which one the channel has ready, and how many older ones the repository still carries. Those older ones are not history: they are where a rollback can go.
The actions on that screen are the ones that already existed: installing or removing a PHP series (removal is refused by the machine while a website still runs on it) and updating the server now. Read channel shows what another channel would offer without moving the machine onto it — and the screen says that list was not checked against this machine's keyring.
Blackout windows
No wave starts on a Friday, Saturday or Sunday, on a Dutch public holiday, or while an incident is open. This is not superstition: a wave bakes for hours and the whole value of a bake is that somebody is watching it.
Trying anyway is refused, with the window named. The only way past it is the emergency lane.
A recovery rollout for what was left behind
A rollout that stops leaves machines behind: the one that failed, and the rest that were never attempted. Inside that rollout nothing happens any more — that is the rule above and it stays.
What you can do is now on the board itself. Under Stopped updates the stopped rollout appears with its reason, and beside it one button: Start a recovery update. It opens a new rollout covering exactly the machines that were left behind — the failed one and the skipped ones, and not the machines that did take the update.
Two things you may meet there:
- An incident is open. A rollout that stopped at a red check opens one itself, and while it is open the panel refuses every new rollout. The board says so. Close the incident, or go through the emergency lane with a reason.
- You picked the machines, so you are the canary choice. In a group at least one machine always stays behind the gate: the canary is a sample. A list you put together is not — you named them one by one — which is why one server is a valid plan here.
The same rollout from a terminal, with the machines in it:
curl -s -b cookies -X POST https://panel1.corecp.dev/api/v1/rollouts \
-H 'Content-Type: application/json' \
-d '{"node_ids":["<id>","<id>"],"canary_count":1,"note":"recovery"}'The emergency lane
The red block on a release's panel. It asks for a reason before it does anything, and it lists what it is about to override:
- the beta soak — hours instead of days,
- the bake between waves — minutes instead of a day,
- a group's channel pin and version ceiling,
- blackout windows,
- the softer health checks.
And what it does not override, which is the part worth trusting:
- the canary. No class may skip it, and there is no path in the panel that plans a rollout without one.
- a red gate. The rollout still halts, still rolls the failing machine back, and still pages.
Everything the lane does is written down: on the release, on the rollout, and in the audit log.
From a terminal
Everything above is the API, and the API is the same one the screen uses.
# What is in the pipeline
$ curl -s -b cookies https://panel1.corecp.dev/api/v1/releases |
python3 -c 'import json,sys
for r in json.load(sys.stdin):
print(r["package"], r["version"], r["risk_class"], r["state"], r["channel"])'
corecp-php85 8.5.3 minor-runtime approved stable
corecp-agent 0.34.1 platform awaiting_approval beta
# Run the promotion policy now instead of waiting a quarter of an hour
$ curl -s -X POST -b cookies https://panel1.corecp.dev/api/v1/releases/tick
# May a wave start?
$ curl -s -b cookies https://panel1.corecp.dev/api/v1/releases/windows |
python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["open"])'
TrueOn a server itself, the same health check the gate uses is one command:
# root@stck1
$ corectl release gateIt exits 0 when the machine is no worse than its baseline and 1 when it is, naming the site or the service that changed.
Everything on this page is in your language
The release board explains a lot: why a window is closed, what a risk class means, who may promote a version, why a soak is not clean yet. Those sentences are composed on the server, and until round 2's final pass they came out in English even when your panel was set to Dutch.
That is fixed, in a way you can see: switch the panel to Nederlands at the top right and the same explanations appear in Dutch, without the page reloading. The server no longer sends a sentence — it sends a reference plus the numbers that belong in it, and the panel builds the sentence.
Nothing changes on the command line, where the English text is still what you get:
ssh panel1.corecp.dev 'corecp-panel releases policy --config /etc/corecp-panel/panel.yaml'
ssh panel1.corecp.dev 'corecp-panel waves board --config /etc/corecp-panel/panel.yaml'If you do still see an English sentence on this screen with the panel set to Dutch, report it: that is a hole in the translation list, not a decision.
See also
docs/maintenance.md— the full policy, the matrix and the mechanics.- Adding a server — a new machine lands in a late wave; this is how it got there in the first place.
When a rollback is refused
The four buttons on a running rollout — cancel, halt, roll back per server and put the channel back — ask for confirmation first. If it is then refused, that dialog stays open with the reason in it. It used to close, and the explanation was on the card underneath, which you had just been covering.
What succeeds still lands on the card: a rollback's per-server report is there, and by then the dialog is gone and the card is what you are looking at.
These buttons also ask you to confirm who you are — a rollout reaches every machine in a wave. That dialog comes up by itself, you type your code, and the operation carries on without you having to choose anything again. Dismiss it and nothing has happened: no rollout, no message, and the button is still sitting there exactly as it was.
A new version that does not show up in a channel
Very occasionally a package is built, signed and added to the repository and the channel still serves the version before it. What happened is that two builds finished in the same second and both tried to create the channel's listing; one of them won and the other stopped there, with everything except that last step already done. The platform now recognises that exact collision and points the channel at the newer snapshot instead of stopping, so it should not happen again — and if it ever does, the fix is to run the build once more. Nothing has to be undone first: the package is already in the repository, and the second run only redoes the step that failed.
You can always see what each channel is serving:
bash scripts/release.sh statusIf a channel is behind the one before it, that line says so and names the command that moves it.
Release sets: a tested set, or one package
A release is a signed set: the exact list of packages that were tested together, with the fingerprint of every one of them, the version tag they came from and the result of the test run that passed on them. A channel can be put on exactly one set — nothing more, nothing less — and one package can be moved on its own without anything else in the channel changing.
What you can count on:
- A set that was changed after it was signed is refused. Change a single fingerprint and the build server will not use the set, even if somebody signed it again: every package also has its own signed manifest, and those have to agree. The channel keeps serving what it served.
- Putting one package on a channel changes only that package. Everything else the channel offers stays exactly as it was, including the older versions that are kept for rolling back.
- Going back to the previous set gives you the previous set exactly — the same list, byte for byte, not a rebuild of it.
- A set is never rewritten. A mistake in a set is fixed with the next version.
From the panel this is a request, like a channel rollback: you name the set or the packages, and the build server carries it out within a minute and reports back. Until then the channel serves what it served; nothing already installed on a server changes.
On the staging repository there are two channels that matter: edge (every build) and beta (the candidates). stable and steady still exist there, but they follow beta: they always serve what beta serves, so a staging server that was installed without naming a channel still gets what staging tests. The production repository is separate and serves stable as a real channel.
From a terminal
On the build server:
# Which sets exist, whether each still verifies, and which channel serves it
bash scripts/release.sh sets
# Put beta on exactly one set
bash scripts/release.sh promote-set beta 0.50.0
# Move one package onto beta, and nothing else
bash scripts/release.sh promote-packages beta corecp-crs=4.25.1+corecp1
# Back to the set beta served before
bash scripts/release.sh rollback-set betaThrough the panel's API (an administrator, with a recent second-factor check):
curl -s -b cookies -X POST https://panel1.corecp.dev/api/v1/release-channels/beta/promote \
-H 'Content-Type: application/json' \
-d '{"set":"0.50.0","reason":"candidate for this week"}'A refusal says what did not agree — for example that a package's fingerprint is not the one its signed manifest names — and names the set or package.
What exactly is in a version
Since this round, two extra files sit beside every published package: a bill of materials listing every component inside it, and a provenance note saying what it was built from and with. Both are signed, and their fingerprints are inside the release manifest you already check.
bash scripts/sign-release.sh verify corecp-agent 0.39.0== attestations
ok corecp-agent_0.39.0_amd64.sbom.cdx.json cyclonedx-1.6 20431 bytes
ok corecp-agent_0.39.0_amd64.provenance.json slsa-provenance-1.0 1685 bytesA release with no bill of materials is not published — signing refuses it. For packages older than this practice the list is empty, so you can tell "there is none" from "nobody looked".
A version without ~dev in its number is, moreover, only built from a protected tag on GitHub; the provenance note names that tag.
The same goes for the packages CoreCP builds from somebody else's source — PHP, Node.js, the firewall rules, imapsync: they are built in a clean environment without internet access and carry both files too. Every package on the edge channel has them. To see for yourself, or after something was published by hand:
ssh root@build.corecp.dev 'python3 /srv/corecp/src/scripts/check-sbom-coverage.py --live' edge: 105 package version(s) in the index, 105 with SBOM + provenance, 0 withoutOlder versions that were published before this practice are taken out of the channel rather than given a bill of materials afterwards. The newest version of a package is never removed that way; it is rebuilt first.
bash scripts/release.sh retention --unattested --dry-run # what would go
bash scripts/release.sh retention --unattested # take them out of the repository
bash scripts/release.sh snapshot # and move edgeBeta, stable and steady keep what they serve until you promote them.
How to fetch the bill of materials for your own version, and check for yourself that it is genuine, is in Where a release comes from.
The production source: rolling out a release
A panel that is the package source of its own servers — the production panel, or the production simulation on staging — shows one more block on this page: Production source. It says what the panel serves (packages.<domain>), where servers fetch the installer (get.<domain>), which build server it fetches releases from, which working key signs today and until when the root key certified it, and which release is rolled out now.
Roll out is one action: enter the version as the build server made it (for example 0.50.0), a reason if you like, and press Roll out. The panel asks for your second factor — this is the act the working key exists for — and then does three things you watch appear in the log box:
- fetching — the release (the signed set) from the build server, with every package, its bill of materials and provenance note, and the log of the test round;
- verifying — the build server's signature on the set and on every package, every file's fingerprint and length, and the gate evidence: the set must record that the production gate passed on it, and the log that came with it must be the log that evidence was made over;
- publishing — the same files, re-signed with this panel's working key, as an apt source. Servers pick it up on their next update; nothing already installed changes.
What you can rely on:
- A set that does not add up is refused, with the reason. A set without gate evidence, gated by another script, with a changed package or with a log that does not belong to the evidence never reaches the source. What the source served stays as it was.
- Back to a previous release is the same button with the previous version: the source keeps the current and three previous releases, so servers can go back too.
- Without a valid working key nothing happens. If the certificate has expired or the key is revoked, the block says so and Roll out is off. What to do then is in the key-management runbook.
- A staging package placed straight into the production source is refused by every production server: they trust only what the working key — certified by the root key — signed.
From a terminal
On the panel:
corecp-panel dist status # what the source serves, which key signs, whether a rollout can run
corecp-panel dist publish 0.50.0 # fetch, verify, publish
corecp-panel dist verify 0.50.0 # verify only (after dist fetch)
corecp-panel dist list # releases rolled out beforeOn a server, corectl update trust shows the source, the root key the server trusts, the working key it saw and until when it is valid.