Your servers' backups
There are two kinds of backup in CoreCP, and they belong to different people. A hosting account's backup is the customer's: they make it, to their own storage, and they put it back themselves (see Your own backups). A server's backup is you
Written for: Administrator
There are two kinds of backup in CoreCP, and they belong to different people. A hosting account's backup is the customer's: they make it, to their own storage, and they put it back themselves (see Your own backups). A server's backup is yours: the whole machine, with /etc, /home, every database and the state CoreCP keeps for itself.
This page is about the second kind. You will find it under Servers → Backups.
A destination first, then servers
Where your fleet's backups go is recorded once, not once per server. Under Servers → Backups → Destinations you record a destination:
- the kind of storage — a rest-server, an sftp server, or S3;
- the address of that storage;
- the path under which every server gets a place of its own;
- the credentials, where the kind needs them.
Every server gets a repository of its own under that path. That is not a detail: restic deduplicates within a repository, and in a shared one every server can reach every other server's snapshots — including to delete them. The password to such a repository is generated on the server itself and never leaves it. The panel only holds the credentials of the storage.
Write the repository password down somewhere off the server. Without it the snapshots cannot be read by anybody, us included. You will find it on the server in /etc/corecp/secrets/backup.env.
Try it before you rely on it
Beside each destination is Validate destination. It first asks which server should do the measuring, and that is deliberate: whether storage is reachable is a property of the route between one machine and that storage. A green tick with no machine attached to it says nothing.
The chosen server then walks a checklist — does the storage answer, is there a repository, does anything actually go in — and writes a small probe snapshot which it removes again. That last part is the point: storage that answers politely and cannot be written to (a full disk, a read-only export, an S3 policy that only allows reads) is exactly what this button exists to catch.
If a step fails, the page names which. In the terminal it is the same measurement:
corectl backup destination check --target sftp \
--endpoint backup@storage.example.net --repo /srv/restic/stck1 --initWhich destination? The panel proposes one
With more than one destination you do not have to remember which has the most room. On the backups page of a server that writes nowhere yet — and in the add-server wizard, once the machine has joined — the panel shows a recommendation, with the measurements it rests on underneath it.
There are five of them, in the order they decide:
- another location. Does the storage stand somewhere other than the server itself? A snapshot beside the machine it protects survives a failed disk, not the building. A destination in the same location is still offered — last, and with a warning beside it.
- what the last check said. A destination whose last validation failed is ranked last and is never the recommendation.
- free space, measured on the storage machine itself.
- how many servers already write there — fewer is better spread.
- the measured distance between this server and that storage.
What the panel has not measured says so. "Free space not measured" counts neither as "plenty of room" nor as "nearly full": it sits exactly between the two. The same goes for a location nobody filled in.
Free space and distance exist only when the destination's address is a machine this panel knows. For storage at a third party the panel cannot know them, and it says so rather than guessing.
The destinations that were not chosen are folded away underneath, each with its reason. One that cannot be used at all — belonging to another server group, with no stored login, or the server itself as its storage — is listed too, in as many words.
Fill in the location of a destination (Servers → Backups → Destinations) and of your servers, in your own words: "fra1", "ams-rack-3". The panel only compares whether two locations are equal; empty means unknown.
When a server runs websites, databases or mailboxes and no destination is recorded for it, the server advisor says so, with a button to that same page.
Attaching a server
Switch the backup role on from the server's configuration panel, or pick a destination at the backup role in the add-server wizard. Attaching does in one go what is otherwise four separate acts: install the role, put the credentials on the machine, write the target into node.yaml and create the repository.
With an sftp destination the server generates a key of its own and prints the public half. Put that in the storage machine's authorized_keys — that is the one act that belongs on the other side.
If you add the server through the wizard the machine does not exist yet, so the choice is recorded and the server's own backup page says afterwards that it still has to be attached, with the button beside it.
The Backups tab is there on a server that takes no backups either. It says what is true and which destination the panel recommends — that server is precisely the one you have to be able to reach. There is no schedule form on it: without storage the server would refuse every change anyway.
Moving to different storage is Move to different storage, above "Where this server writes". Mind what it does and does not do: the snapshots in the old repository stay there — they do not move with it and they are not cleaned up — and the first run to the new storage is a full one, because there is no history there yet to build on.
Schedule, retention and pruning
Servers → server → Backups shows three runs in a row, in the order the night runs them: the server, then the accounts, then the retention pass. They are staggered because the account pass reads what the server pass has just walked.
Each run is picked as how often plus at what time — "every night at 01:00", "every week on Sunday at 04:00". Behind Advanced is the calendar expression itself; that is literally what the server is given, and you only need it for schedules the controls cannot say (twice a night, weekdays only, a list of dates). Under each run is when it fires next, out of the server's own answer.
A schedule the server cannot read is refused, with the server's own sentence under the box it was about. Nothing is saved then — your typo cannot stop a timer. If a timer is stopped anyway (a server from before this version, or a hand-edited file), the page says so in words instead of drawing it green.
Do they all start at once?
Every attach takes the default time, so a fleet converges on one minute by itself — and then ten servers write to the same storage at once. If this server collides with neighbours on the same destination, the screen says so: which machines, on which storage, which slot is the quietest in that night, and how many minutes of room it has around it. One button fills it in; you press save.
What the panel cannot read as a plain clock time is not counted, and it says so. An advice that quietly treated those schedules as free minutes would recommend a time it never checked.
How long is a backup kept?
Retention is four groups — daily, weekly, monthly, and a floor regardless of date. They apply per group: "7 daily" means seven days for each account and for the server layer, not seven snapshots in total. And note what a number in a group actually means: "keep daily 7" keeps the last snapshot of each of the last seven days, not the last seven snapshots. The screen says that under each box, so you do not have to remember it.
Each group has a minus and a plus, and stepping below 1 lands on none — the word, not a number. That is deliberate: "keep none of this group" used to be a minus sign five pixels wide next to a 1, and the two mean opposite things. The floor is the odd one out: its "off" is none as well, and it simply means there is no floor — and setting it to none really removes it, on a pattern and on a single server's own form alike.
You cannot turn all four off. A policy that keeps nothing is refused by the server, with the group it is about named — the backup tool refuses such a policy outright, so the retention pass would fail every week without removing anything, and nothing on the screen would say so.
One rule for the whole fleet
Typing four numbers per server is how two servers drift apart with nothing to say so. Under Servers → Backups → Retention patterns you write the rule down once, give it a name, and point servers at it.
Changing a pattern is two presses, and the first one is the important one. Save records the rule and touches no server. Apply is the second press, and before it happens the panel says what it is about to do:
- how many servers end up with different numbers,
- how many of them will keep less than they keep now — the direction in which snapshots get deleted,
- and how many the panel has never read a retention from, which is neither of the two above and is counted on its own.
If one server is unreachable, the others still get the rule. The report says which one refused and what it said, and one button retries exactly that server. Nothing is rolled back: a server that took the numbers has them, and undoing that because a different server failed would be a second change nobody asked for.
On a server's own page you can see which pattern it follows and whether it has actually taken the numbers. "Recorded, but not on this server yet" is its own answer — it is the honest one for a machine that has been unreachable since you put it on the rule.
Is the cleanup actually running?
The retention pass can fail quietly. The most common way: it takes an exclusive lock on the repository, and a run that was killed while holding one leaves that lock behind — every later pass then fails against it, while the backups themselves keep working and the timers keep looking healthy.
So the server's backup page says whether the last cleanup pass finished, never ran, or failed, and with what. "Nobody has run one yet" is its own answer and never counted as healthy. The pass now clears a lock the backup tool itself considers abandoned and tries again — before its very first step, because such a lock refuses even a read of the repository; a lock a running pass is holding is left exactly where it is.
The note on that page appears whenever there is something to say, and that includes the two findings that are easy to miss: a server that keeps nothing at all, and storage that is running out of room. When there is nothing to say, the page stays quiet.
If your storage refuses deletes (a rest-server started with --append-only), pruning from this server removes nothing. That is not a fault but the design: writing and erasing deliberately do not live on the same machine. The screen says so, and the retention pass then belongs on the storage machine.
Restoring
Restore… lets you pick a snapshot. By default it lands in a separate directory and the running system is untouched — which is what you want when you only need to see what was in /etc last week.
To write over the running system you have to type the server's name out. The panel does not fill it in for you: a confirmation the panel could produce is not a confirmation. Run a reconcile on that server afterwards, because the state underneath the services has changed.
Who is behind?
The list under Servers → Backups is ordered by how far behind a machine is: one that has never produced a snapshot first, then the stalest. A server with no backup role is on the list too, saying so in as many words — precisely because "it is not in the list" reads as "nothing is wrong".
Behind is the same line the server itself uses: corectl doctor fails a machine whose newest snapshot is older than two days. One missed night is a hiccup, two is a pattern.
Is it ever tested that they can be read back?
Yes. Once a week every server with the backup role takes a sample out of its newest snapshot on its own, reads it back and throws it away again. The server's backup screen says when that last went well; when it went wrong it turns up in the list of problems, with the server's own sentence beside it. A server that has snapshots and has never had one read back is a promise — this turns it into a fact.
You can also run it yourself with Run the restore test now, and switch it off with corectl backup configure --restoretest off.
A backup server of your own that lets nothing be deleted
The safest destination is a CoreCP backup server: a server of your own (b1 in production) that keeps your other servers' backups and refuses to ever delete one. Whoever takes over one of your servers can write new snapshots with it, but cannot erase old ones or read another server's.
- Set the backup server up once from the terminal (see below). After that it is an ordinary server in the panel.
- Under Servers → Backups → Destinations → Add destination, choose the kind CoreCP backup server and fill in its address, for example
https://b1.example.net:8000. You do not fill in a login: when a server is attached, the backup server gives it a login of its own. - Attach your servers as usual. The panel asks the backup server for a login and a key and puts them on the server; the panel itself keeps neither. This needs full access to both servers.
Cleanup happens on the backup server. An attached server can no longer remove anything itself, and Clean up on its page says so. Every night at 05:30 the backup server applies the retention you set on the server's page — a change there is passed on by itself. It counts in time windows ("everything from the last 7 days, one per day") rather than in numbers, so fake snapshots cannot push the real ones out. A policy with only "keep the last N" is refused for that reason.
The Immutable check in the server's health is red for this kind when the backup server would allow a delete after all.
From the terminal
On the backup server:
corectl backup target serve # rest-server, append-only, TLS
corectl backup client add w1 --from 203.0.113.10 # login and key for w1, shown once
corectl backup client set w1 --keep-daily 14 # w1's retention
corectl backup target prune --dry-run # what the cleanup would remove
corectl backup target statusOn the server that writes to it (the login and the key are read from stdin, never as an argument):
corectl role add backup --target rest-append-only --endpoint https://b1.example.net:8000 --repo w1
corectl backup configure --ca-cert /etc/corecp/agent/ca.crt
corectl backup credential set --rest-user w1 --rest-password-stdin
corectl backup init
corectl backup target-key add --password-stdinThe backup server also keeps an eye on the panel and mails when it has not answered for longer than the time you set:
corectl backup target watch --url https://p1.example.net --after 10m \
--mail-to ops@example.net --relay-host mail.example.net --testThe panel's own database goes to the same server with pgBackRest; how to set that up is in the panel's recovery plan.
See also
- Restoring a backup — the journey from snapshot to verification.
- Your own backups — the backup a customer makes and restores themselves.
- Adding a server — where the destination sits in the wizard.