THE SELF-HEALING LAB

SRE discipline for a one-person fleet — incidents that never page a human

Site Reliability EngineerVirtualization / systemd / watchdogs / LLM triage2025 — NOW

The Problem

Traditional SRE wisdom: page someone when something breaks. This fleet has no someone to page. A small self-hosted cluster — a hypervisor, a low-power sensor node, container hosts, and local GPU inference — runs three always-on AI daemons that execute work autonomously, 24 hours a day, with no on-call rotation and no ops team.

The design question was not "how do we page faster" but "how do we stop paging at all." That meant building a system that could detect its own failures, classify them, attempt self-recovery, and — only when it genuinely couldn't proceed — hand a diagnosis to the operator rather than a raw alert.

The goal isn't faster incident response. It's incidents that never reach a human.

Detection Architecture

The alert stack is deliberately over-specified for the scale. Determinism handles the easy cases. A local large language model handles the hard ones. Neither trusts the other blindly, and neither makes a network call to an external provider.

NETWORK IDSsensor nodenetwork tapPRE-FILTERrule-based triagedeterministicLOCAL LLM VETlocal GPU inferenceno external API callinjection guard activePUSH ALERTend-to-end verified ✓synthetic injection 2026-07-131. CAPTURE2. TRIAGE3. CLASSIFY4. NOTIFY⚠ current verdict sample is small — pipeline validated, rates not yet characterized

Why local inference. The LLM classifier runs on local GPU inference (a local open-weights model at interactive speed). No external API call, no data egress, no per-alert cost, no rate limits from an outside provider. The tradeoff is compute weight for a classification task that a smaller model might handle. The payoff is zero keys at risk, deterministic testability, and a classifier that can be interrogated locally before being trusted with live alerts.

The small-sample caveat. The current sample of vetted IDS verdicts is small. That is enough to validate the pipeline path, not enough to characterize false-positive or false-negative rates — both of which are labeled UNMEASURED. The alert path itself was verified end-to-end with a synthetic injection on 2026-07-13: a real event travelling from the IDS through the local classifier to push delivery, confirmed at the receiving endpoint.

Self-Healing Patterns

Three patterns compose to keep the system alive without human involvement. None are novel in isolation. What matters is how tightly they are specified and how faithfully they are documented when they fail.

Pattern 01 — Watchdog FSMs for daemon liveness

Each daemon runs as a supervised service with a watchdog that classifies its state into a finite set of transitions: idle → active → frozen → restarted. The FSM prevents restart-loops on transient states — "activating" and "heads-down conductor" are distinct from "frozen" and do not trigger recovery.

A critical production lesson: a liveness detector keyed on CPU and I/O counters reads GREEN on a reply-side wedge, because idle polling advances those counters even when the agent loop is frozen. The FSM distinguishes reply-delivery liveness from process liveness. Both signals are required; neither is sufficient alone.

Pattern 02 — Session-state resume after restarts

Each daemon owns a per-lane state file recording what it is executing right now and the next planned step. The file is written at step boundaries and cleared on completion — not every turn. On restart, the boot brief surfaces the state file first, so a restarted daemon resumes warm rather than cold-starting and dropping in-flight work.

The failure mode this closes: a daemon that hits a session limit mid-turn never auto-resumes after the window resets. Without the state file, queued work rots silently while every watchdog reads green — the worst failure mode because it has no alarm. With it, restart equals resume.

Pattern 03 — Guaranteed-delivery alert backstop

The alert path has a backstop that verifies delivery at the destination endpoint, not that the sending service is merely active. That distinction matters: a service can report "active," dispatch a payload, and still silently mis-route the alert to the wrong host or stream.

Verification method: synthetic end-to-end injection (performed 2026-07-13). A synthetic event traverses the entire pipeline — IDS tap through LLM classifier through push delivery — and is confirmed at the receiving device, not inferred from server-side logs. "Service active" is not the same as "alerts delivered."

Liveness is not "process alive." It is "the right output reached the right destination in the last N seconds." Every other definition is an optimistic proxy.

Incident Postmortems

Three real incidents from the permanent archive. Each produced an entry in an append-only lessons corpus — 160 entries spanning March–July 2026 — so the same failure cannot recur without a prior warning. The format is the differentiator: every incident resolves to a headline rule that the next session inherits before acting.

PM-01A hypervisor firewall toggle silently blackholed all VM traffic
Timeline
Enabled a hypervisor-level firewall. All off-host VM traffic stopped. Host-local traffic and health checks continued normally. No immediate error surfaced.
Root cause
Turning on the hypervisor firewall changed a kernel networking setting as a side effect, and the hypervisor firewall then silently dropped bridged traffic. Turning the firewall back off did not undo the side effect — the change persisted across the toggle.
Fix
Rolled back the firewall change, verified the underlying kernel state directly rather than trusting the toggle, and restored it. Traffic recovered immediately.
Detection gap
Standard host-local liveness checks never traverse the bridge. Every health probe continued to pass throughout the outage. Cross-bridge probes are now required.
Lesson L089 — permanent archiveA toggle that changes hidden system state is not reversed by toggling it back. Verify the underlying state, not the switch — and probe across the network boundary, because host-local checks stay green while every off-host packet is dropped.
PM-02Host backup freeze caused fleet-wide false “network down” alarms
Timeline
A scheduled host backup froze the container filesystem mid-operation, halting all daemon activity. On resume, the liveness watchdog observed stale probe timestamps and fired “chat / network broken” alerts across multiple daemons simultaneously.
Root cause
The watchdog’s stale-probe detector did not distinguish between a genuine network outage and a container-level wall-clock jump caused by a backup unpause. Uptime and wall-clock advanced differently during the freeze — the delta is the discriminating signal.
Fix
Added a suppression condition: if the watchdog detects a wall-clock jump (not just an uptime delta), suppress the stale-probe alert for one grace period. Backup resume is the signal to suppress — not the command channel timestamp.
Detection gap
The watchdog treated uptime and wall-clock as equivalent. A freeze-then-unpause advances wall-clock without advancing uptime — that difference was the root cause, not a network failure.
Lesson L094b — permanent archiveA hung host backup freezes the whole container; on resume, the stale-probe watchdog false-alarms “chat/network broken.” Suppress after unpause (wall-clock jump), not just after boot (uptime). One broken lane is not darkness — require two independent signals before declaring an outage.
PM-03An IDS suppression rule silently failed to load
Timeline
An IDS suppression rule was added to mute a noisy alert. No error was raised. The suppressed alert kept firing. The mute appeared effective from the config alone.
Root cause
A formatting quirk in the rule file caused the parser to skip the directive without surfacing an error. The rule was present on disk but never applied, so the alert fired as if no suppression existed.
Fix
Corrected the rule formatting and verified the mute by watching the alert stop — not by verifying the config file contained the directive.
Detection gap
Config-level verification (“the rule is there”) is not behavioral verification (“the rule is applied”). The only true test is observing the suppressed alert stop arriving.
Lesson L123 — permanent archiveA config rule that fails to parse can fail SILENTLY: the directive is on disk, nothing loudly reports it, and the behavior never changes. Verify a mute by watching the alert STOP, never by the config containing the line.

Stack

KVM virtualizationsystemdNetwork IDSLocal open-weights LLMWatchdog FSMsLocal GPU inferenceLow-power sensor node

Result

What "healthy" means on this fleet is measured, not asserted. The numbers below come from the verified corpus — not from a monitoring dashboard that reports uptime without defining what it is counting.

Measured & verified
160append-only engineering lessons in permanent archive (Mar–Jul 2026)
1,564work receipts on disk (1,002 active + 562 archived)
406persistent memory files; 121-entry curated index
3always-on daemons (restart-safe, session-state resume)
50installed skills; 21 learned beliefs (6 gates, 15 advisory)
1synthetic end-to-end alert injection, delivery confirmed 2026-07-13
27findings from independent adversarial Constitution red-team, all addressed
UNMEASURED — labeled honestly
—MTTR before and after system introduction — no baseline exists
—Uptime % per host — monitored but not aggregated or trended
—IDS false-positive rate — verdict sample too small to characterize
—IDS false-negative rate — no ground-truth corpus for this traffic
—Cost savings in $ — no pre-system labor baseline to compare against
—Daemon restart frequency — watchdog fires not yet aggregated into a trend
—Session-state resume success rate — feature too new; sample too small

The UNMEASURED column is not a failure — it is an honest instrumentation backlog. A metric labeled UNMEASURED tells you exactly where to invest next. A metric invented to fill the gap tells you nothing and erodes trust in everything beside it.

The archive holds 160 lessons in the postmortem format shown above, append-only, indexed by ID with a one-line headline rule. The index is regenerated on a schedule. Every correction becomes a rule the next session inherits before acting. Institutional memory does not expire when a model does.

Radical honesty beats managed optics. Missing data is more useful than invented data.