Server health
Agentless server monitoring over SSH: resource hogs, disk-full and memory-leak forecasts, failed units and OOM kills, top processes, and Farabi insights.
The metric strip shows numbers. Health says what they mean: which process is eating the CPU, which disk fills up and when, what is unusual for this hour, and what failed since you last looked.
Nothing is installed on the server. Health reads the metrics the terminal already samples, plus a few read-only commands over the SSH session you opened. Every signal is worked out on your computer, by fixed rules, with no model involved. A model never decides whether something is a signal, how bad it is, or when something runs out.
Detection is free
Signals, the Health panel, the top processes, snooze, acknowledge and notifications are part of the Free plan. Farabi’s explanation of them is part of Gatesys Pro. See Plans.
The Health pill
A pill beside the metric strip in the terminal header sums up the host: Healthy, or a count such as 3 issues · 1 risk.
- An issue is happening now: a process hogging the CPU, a failed unit.
- A risk has not happened yet: a disk that will fill in three hours, a kernel update waiting for a reboot.
The pill takes the colour of the worst signal. A metric card with a signal on it wears a small dot in the same colour, and a critical dot breathes. In the Hosts list, each host carries a badge with its worst colour and its count of signals.
Click the pill to open the Health panel. Esc or a click outside closes it.
What Health watches for
| Kind | What raises it |
|---|---|
| Resource hogs | A process at 80% of a core or more in two samples in a row, while host CPU holds at 65% or more for a minute. Critical when host CPU stays above 90%. Memory works the same way: host memory at 80%, and a process holding a fifth of it. At most three per kind, the worst first |
| Forecasts | A straight-line fit over the last hour says when a disk fills or runs out of inodes, such as /var full in ~3 h. Raised when the fill is due within 24 hours, critical within one. A filesystem at 90% is a signal whatever its trend, critical at 97% |
| Memory leaks | A process whose resident memory climbs in a straight line (six samples over ten minutes, a close fit, at least 5% growth), with its growth per hour and when memory runs out |
| Saturation | Load above the core count for two minutes (critical past twice the cores). CPUs waiting on disk (I/O wait) 20% of the time for a minute (critical at 40%). Swap at 10% and growing, or at 50% (critical at 80%) |
| Unusual for this hour | CPU, memory or load well above what this host usually shows at this hour of the day, such as CPU 71% — usually 8–22% at this hour |
| Events | From the host facts: an out-of-memory kill, a failed systemd unit, a container restarting or unhealthy |
| Pending risks | A kernel updated but not rebooted, a growing number of zombie processes, the server clock 30 seconds or more off, a certificate close to expiry where a fact carries one |
A forecast needs enough points, over enough time, on a line that fits. With too little data, no ETA is shown rather than a guess.
The normal for each hour is learned per host from the moment Gatesys SSH starts watching it, and kept for a week, sealed on this computer. The panel says how long it has been watching.
Signals do not flap
A signal is raised only after two checks over its line, holds inside a lower line, and clears only after three checks without it. It gets worse at once, and eases only once the lower severity has held. The same thing is always the same signal, updated in place, never listed twice.
Read a signal card
The panel lists signals worst first. Each card shows:
- Severity and kind: Critical, Warning or Note, and Process, Forecast, Saturation, Unusual, Event or Pending.
- How long it has held, and whether it is rising or falling.
- The title, such as java pid 2211 at 340% CPU for 6 min.
- What happens if you ignore it, and when: When it fills, writes fail: logs stop, databases refuse writes and logins can break — in ~3 h at this rate.
- A line of the last hour, with the forecast drawn past it as a dashed line inside its band.
- Evidence chips: the measured values behind the signal. Hover one to see the read-only command that shows the same thing.
- What it is tied to: a measured change in the half hour before it began, such as Started 2 min after “nginx 1.24.0 → 1.25.3”, and, from the service map, who uses the ports involved, such as 3 known hosts held connections to :5432.
- The processes behind it, when there is more than one.
Act on a signal
| Action | What it does |
|---|---|
| Look closer | Opens read-only commands for this signal, such as the biggest folders on a filling disk. Pick one, then Insert only or Run |
| Ask Farabi | Asks Farabi to explain the signal from the evidence. See Ask Farabi about health |
| Watch it | Opens the Watch composer with a metric watch filled in, such as disk above 95% for two minutes. See Watchers |
| Open as Safe Change | Where the fix is a config file: journald’s size cap for a disk filling with logs, vm.swappiness for swap, timesyncd for a drifting clock. See Safe Change |
| Snooze | Hides the signal for 1 hour, 8 hours or 1 day |
| Acknowledge | Hides the signal until it gets worse. If it clears and later comes back, it is news again |
Snoozed and acknowledged signals are counted at the foot of the list: 2 snoozed or acknowledged · Show again brings them back. Both are kept per host.
Top processes
Under the signals, Top processes lists what the last sample saw: each process’s name, PID and user, CPU, memory (MEM), resident memory (RSS) and age. Click a column header to sort by CPU, memory, RSS or age; click it again to reverse the order.
- CPU is per core: 100% is one core. The header says how many cores the host has.
- Values past the rules’ lines are marked, and a process that is a zombie or waiting on disk says so.
- While the panel is open, the table refreshes every 10 seconds and each row draws its own CPU line.
- Each row has Look closer and Ask Farabi.
The process sample is one read-only ps command, at most every 10 seconds, taken only while CPU or memory is hot or the panel is open. Once a minute, a df reads the use of every local filesystem for the fill forecasts. Both run on the session’s background channel, only when it is free: never while a shell is opening, never queued behind another job. On a server that refuses a second channel, sampling stops and the card says why.
Ask Farabi about health
Pro
Farabi’s explanation of health is part of Gatesys Pro — free for 3 months, then $20 a year. See plans
You can ask Farabi about a metric, a signal or a process:
- A metric: hover or focus a card in the metric strip, or a stat tile on the host’s dashboard in Hosts, and click Ask Farabi. The Look closer dialog for a metric has it too.
- A signal: Ask Farabi on its card.
- A process: Ask Farabi on its row in Top processes.
Farabi’s panel opens if it is folded and asks at once, with a question such as Analyse CPU on db-primary, in Farabi’s language. With the question goes a Health block built by the app: the metric’s last hour as minimum, average, maximum, 95th percentile and slope, the active signals and their evidence, the top processes, related facts and changes, and who uses the ports involved. Each line has an id such as [e3].
The block goes through the same path as any question: fenced as data, checked by the Injection Shield, secrets masked, and, for a model off this computer, addresses removed and other hosts replaced with placeholders. What Farabi saw in the thread shows the block exactly as it was sent. A matching skill, such as Linux triage or disk cleanup, is suggested.
The Insights card
Farabi answers with an Insights card:
- A short summary.
- Numbered claims, each citing the evidence it rests on, such as
e3. Hover a citation to see the measured line. - Recommendations, each with why, a risk (Low, Medium or High risk) and its commands.
The app checks every part of the answer. A claim that cites evidence the Health block never listed is dropped whole, and the card says how many, such as 1 dropped. A recommendation’s risk is never lower than the app’s own rating of its commands: the model can raise it, never lower it. Each command is a normal command card, with Insert, Run and their confirms.
A model that cannot answer in that shape still gets its answer shown, as plain text, with the app’s own list of signals under it: What GateSys measured.
Without Pro, or with the assistant off
Ask Farabi then opens the Health panel on the signals instead. Without Pro, a card says the signals stay free and Farabi explains and ranks them with Pro. With the assistant off in Settings › Farabi, a note says the signals are the app’s own and need no model.
Notifications and hooks
- A critical signal due within the hour, such as a disk that will be full in 40 minutes, sends an OS notification when you are away from the window. It uses the same switch and limits as watch alerts: Settings › General › Notifications › System notifications.
- Every raise, escalation and clear is a
health.signalhook event.
On this computer
Health works in a local terminal too, from this computer’s own CPU, memory, load, disk, swap and process table. Open as Safe Change is not offered there.
When Health has nothing to say
- Metrics polling is off. Health reads the metrics samples, so with Settings › Hosts › Monitoring › Metrics polling at
0it has nothing to work from. - Events and pending risks need host facts. OOM kills, failed units, containers, pending reboots and clock skew come from the host facts read. With Remember host facts off, those signals do not appear.
- Nothing needs attention. The panel says so, and still shows the top processes.
Something unclear or wrong? Tell us.