@@PRODUCT@@

If the panel goes down

This article is about the worst case: the machine the panel runs on is gone. A fire, a deleted VM, a disk that does not come back.

Written for: Administrator

This article is about the worst case: the machine the panel runs on is gone. A fire, a deleted VM, a disk that does not come back.

The first thing to know is the reassuring part.

Nothing of your customers' goes offline

If the panel disappears, the servers carry on doing what they were doing. Websites are served, mail is delivered, DNS is answered, cron runs, certificates renew. The agent on each server keeps working from the configuration it already has; it does not need the panel in order to keep working, and there is no situation in which the absence of a panel turns anything off or throws anything away.

What does stop is management: nobody can create an account, change DNS, start a rollout or sign in. Statistics are still collected on every server and kept for up to 35 days; as soon as the panel is back they are pulled in.

So you have time. There is no meter running on customer downtime.

What we promise, and what we measure for it

promisewhat it means
Recovery time4 hoursfrom "the panel host is gone" to an administrator signing in again
Recovery point5 minutesat worst you lose the last five minutes of panel changes — no websites, no mail, no DNS

Those two numbers are not assumptions. Every month the backup is really restored into a throwaway database, and that drill puts a clock on itself: how old the newest recovery point is, and how long restoring took. If the drill does not pass, a mail goes out. If the drill stops running, a mail goes out too — a control that is silent when it is healthy cannot be told from one that is silent because it is dead, so there is a separate watchdog on it.

The four things you have to arrange yourself

Recovery only works if these four exist before you need them. None of them is something the panel can do for you: they are acts of a person.

  1. Two printed panel kits. The kit carries the key every secret in the database is sealed under. A backup without that key gives you back every account, domain and the audit log — and every sealed column stays unreadable. Print two, keep them in two different places.
  2. Two printed build kits. These carry the key updates are signed with, the key of our own certificate authority, and the password of the backup repository. Two copies again, in two other places.
  3. No safe holding both kinds of kit. The whole separation is that whoever holds the backups cannot read them, and whoever holds a key holds no data. One drawer with everything in it undoes that.
  4. A second copy of the backups. It is pulled every night to a second machine, encrypted — that machine cannot read the contents and does not need to.

Want to know whether all of that is still true? The drill prints the human half as a checklist:

ssh root@build.corecp.dev corecp-restore-drill --checklist

How to see that it is in order

ssh root@build.corecp.dev corecp-restore-drill --result
ssh root@build.corecp.dev corecp-drill-deadman --check

The first shows the last drill with its clock. The second says whether that drill is still actually being run.

And if it really happens

Then do not follow this article but the runbook: docs/dr-runbook.md in the source. It is written for somebody woken at three in the morning, it is step by step with every command in it, and it is walked again against the real machines every round.

Two places in the panel appear in it, and you will find them here:

  • taking a server out of service is on that server's own page, under Danger zone;
  • what the fleet refuses is under Security → Withdrawn certificates.

See also

  • Checking, exporting and monitoring the audit log
  • The backups of your servers