@@PRODUCT@@

What the panel knows about itself

The panel watches your servers. Every machine runs an agent that keeps an eye on its own disks, services and certificates, and the fleet screen rolls all of that up into one list.

Written for: Administrator

The panel watches your servers. Every machine runs an agent that keeps an eye on its own disks, services and certificates, and the fleet screen rolls all of that up into one list.

One machine was never in that list: the one the panel itself runs on. No agent runs there. Whatever goes wrong, goes wrong to the tool you would repair everything else with.

This article is about the four things the panel now knows about its own machine, where to read them, and what to do when one of them is red.

Where to find it

Servers → Monitoring, at the bottom, under This panel.

When everything is in order it is four short facts and nothing else. That is deliberate: a green banner that is always there is a banner people stop reading. The moment one of the four has something to say, the finding appears above them as an ordinary attention row — the same shape as every other finding in the panel, with the same button beside it.

Under the block there is one line: when the panel last measured itself. The screen never measures — it reads the last round, which runs by itself every quarter of an hour. A screen left open on a wall should not cause measurements.

Who may see it: administrators. All four are about the control plane itself, and there is nothing to be done about any of them without access to the machine.

The four things

The fleet authority

This is the certificate every server was joined against. The panel uses it to check that a machine is who it says it is.

When it expires, no node certificate can be renewed and no new join can be verified. Replacing it means re-joining every machine — a planned piece of work with a change window, not an afternoon's maintenance.

So the panel warns about it a long way out: amber a year ahead, red three months ahead. That is far longer than for an ordinary certificate, and that is exactly the point.

Do not confuse it with the fleet certificate — that is the certificate the panel identifies itself to the servers with, it renews itself, and it has its own line on the security page.

The panel's own disk

If the panel machine's disk fills up, the screens, the rollouts and the audit log stop at the same moment — and clearing space afterwards costs more time than clearing it now.

There is one row per filesystem, with a bar beside it. What is measured is what an ordinary writer can have: the space the system reserves for root does not count, because the panel does not run as root.

The row looks at two things at once: space and files. The second is the failure nobody sees coming — a disk that looks half empty and still says "no space" because the inode table is full.

Amber at 90%, red at 95%. The same line as for the servers, because a filesystem is a filesystem.

This is one of the two findings that wake somebody.

How far back the panel can be restored

The panel continuously forwards its database to another machine. How far back in time a restore can reach depends on how fresh that forward is: when it stalls, the panel keeps working and only the newest moment it could be restored to stops moving.

That is the nastiest kind of failure there is — nothing visible changes, and you find out on the day you need it.

So the panel reads three things: is the forwarding switched on, how old is the newest forwarded piece, and has the forwarding refused anything recently. The last of those is the earliest signal — it means the forward has already stopped while the age has not caught up yet.

The measure comes from the database's own setting rather than from a number in the panel: a panel tuned to a longer interval is judged against its own.

This is the second finding that wakes somebody.

The last restore drill

Every month the panel's database is actually restored somewhere else, started and questioned: is this really a panel? That drill runs on another machine, because a control a platform runs on itself proves nothing about that platform being gone.

The drill is silent when it passes. Silent is also what it looks like when it has stopped running, and that is why this row exists: it shows the receipt of the last drill, not the silence.

The row goes red when no drill has ever reported, when the last one failed, or when no receipt has arrived for more than forty days. It goes amber when the drill passed but missed one of its own objectives — for example when a restore would cost more than five minutes of data.

What to do when one is red

RowWhat to do
Fleet authorityPlan the replacement. This is work with a window: every machine is re-joined.
The panel's own diskFree space on the panel machine. Usually log files or old backups.
Restore pointLook on the panel machine at why the forwarding is stalling — usually the storage at the other end is unreachable.
Restore drillLook on the machine that runs the drill at what failed. If nothing has arrived for weeks, the drill is no longer running.

All four also reach you through the bell, push notifications and Telegram, according to your own preferences under Servers. The disk and the restore point additionally wake somebody; the authority and the drill do not — you read those on a screen in the morning, which is what screens are for.

"Not measured" is not "in order"

If the panel cannot make one of its three own measurements, that row says not measured, with the machine's own error beside it. It then goes amber and never green.

That is deliberate: a measurement that did not happen is not a measurement that passed. Drawing green there would make "I could not look" look exactly like "nothing is wrong".

On the machine itself

If something is red and you are already at a terminal on the panel machine, this gives the same four rows the screen does — with the measurement under them:

corecp-panel selfwatch --all --config /etc/corecp-panel/panel.yaml
this panel: measured 2026-09-08T01:12:44Z
SEVERITY  CHECK                          DETAIL
OK        platform.security.fleet-ca     CoreCP Root · 2035-08-11 · 3259d · 1x
OK        platform.perf.panel-disk       2x fs · max 41% bytes · max 12% files
OK        platform.config.wal-archive    last_archived 3m ago · objective 5m
OK        platform.config.restore-drill  ok · 10 checks · rpo 170s/300s · 2026-09-07 · 24h ago · build.corecp.dev

To measure again instead of waiting for the quarter-hour round:

corecp-panel selfwatch refresh --config /etc/corecp-panel/panel.yaml

The command exits non-zero as soon as one of the four is red, so a timer or a check script can branch on it without parsing anything. An amber row is not that: it is on the screen, and it is not a failure of this command.

See also

  • If the panel goes down — what happens when the machine really is gone, and how the recovery runs
  • Email we send you — which findings travel through which channel, and how to set that