Operating-system maintenance and maintenance mode
Your servers install their security updates every night. Two things that go with that are not a server's to arrange: that the alarms are quiet while you work on it, and the order the fleet goes down in when something has to reboot. That is
Written for: Administrator
Your servers install their security updates every night. Two things that go with that are not a server's to arrange: that the alarms are quiet while you work on it, and the order the fleet goes down in when something has to reboot. That is what this article is about.
Putting a server into maintenance mode
Go to Servers → Maintenance, or open the server itself. Putting a server into maintenance mode does three things:
- a restore point is taken first — a snapshot of the server's files plus a dump of its databases. It is the same restore point the nightly backup makes, only now;
- the alarms about that server go quiet. If it drops off, nobody is woken and no push message goes out. The log still records it;
- no new customers are placed on it. Everything already there keeps being served — exactly as on a server you are emptying.
You always give a reason. It is what you (or a colleague) read back in a week when you wonder why that server was quiet at night.
If the restore point cannot be taken, maintenance mode does not go on. Going quiet with no safety net is worse than neither: fix it first, and only force it when you are knowingly taking that risk. A server with no backup destination at all is a different case — there is nothing to fail, it is said out loud, and the mode goes on.
Taking it back out
When you take the server out of maintenance mode it checks itself: the full doctor run plus a real request to every website on that machine. The verdict is in the answer.
The server always comes back out, even when that check finds something. That is deliberate: a server that stays quiet because a check hung is exactly the outage nobody notices.
If the server happens to be unreachable at that moment, the panel puts its own records straight — the alarms are live again — and tells you the machine itself still thinks it is in maintenance. Repeat it once it answers, or do it on the machine.
Services still running old code
An update replaces a shared library on disk. Programs that already had the old one open keep running with it. Your server can therefore be fully updated and still be serving the hole you just closed.
CoreCP asks needrestart which services those are and restarts precisely those. Three never: the agent, the panel and ssh — the three ways of losing what you are working with. Those are reported as deferred, with the reason, so you can see what is still open.
If a newer kernel has been installed, that is reported separately. A restarted service does not fix a kernel.
The reboot plan
Servers → Maintenance carries the reboot plan at the bottom: which servers go down together, and in what order. Everything in one slot goes at the same time; the slots run in order. A slot can fall on a later night — that is how you write down that one group goes a day before another.
The panel refuses three things, with the reason:
- every nameserver of a set at once. That set answers for the domains hanging off it;
- a zone's primary together with its own secondary. Then there is neither a source nor a copy covering for it;
- the machine this panel runs on, unless you approve it explicitly. For the length of that reboot there is no panel — including for the servers you took down just before it.
You get all the points at once, not one at a time. And the approval belongs to this plan: take the machine out of the plan and the tick goes with it.
Updates in waves
Staging and production both run at three in the morning. The difference is not the hour but what they are offered: the mirror puts today's snapshot in front of staging, and production moves onto it only once it is a day old. A bad package has then spent a day on an environment with no customers before it reaches the ones with.
There is nothing to configure. What you do need to know: a fix you see on staging today reaches production tomorrow.
The summary by mail
The morning after a maintenance round the alert address gets one message with, per server, what happened: how many packages, how many services restarted, and whether the server is updated, rebooted, still waiting or failed.
Every server is in it, including the ones with nothing to do — otherwise you cannot tell "nothing to do" from "never asked". A server that could not be reached is in it too, with that outcome.
From the terminal
root@p1:~# corecp-panel maintenance mode on w1.example.net --reason "kernel 7.0.0-31"
root@p1:~# corecp-panel maintenance mode
root@p1:~# corecp-panel maintenance mode off w1.example.net
root@p1:~# corecp-panel maintenance reboot-plan
root@p1:~# corecp-panel maintenance reboot-plan --check /tmp/plan.json
root@p1:~# corecp-panel maintenance summary --sendOn the server itself:
root@w1:~# corectl node maintenance-mode status
root@w1:~# corectl node restart-services --dry-run
root@w1:~# corectl node restart-services