THE SELF-HEALING LAB
SRE discipline for a one-person fleet — incidents that never page a human
The Problem
Traditional SRE wisdom: page someone when something breaks. This fleet has no someone to page. A small self-hosted cluster — a hypervisor, a low-power sensor node, container hosts, and local GPU inference — runs three always-on AI daemons that execute work autonomously, 24 hours a day, with no on-call rotation and no ops team.
The design question was not "how do we page faster" but "how do we stop paging at all." That meant building a system that could detect its own failures, classify them, attempt self-recovery, and — only when it genuinely couldn't proceed — hand a diagnosis to the operator rather than a raw alert.
The goal isn't faster incident response. It's incidents that never reach a human.
Detection Architecture
The alert stack is deliberately over-specified for the scale. Determinism handles the easy cases. A local large language model handles the hard ones. Neither trusts the other blindly, and neither makes a network call to an external provider.
Why local inference. The LLM classifier runs on local GPU inference (a local open-weights model at interactive speed). No external API call, no data egress, no per-alert cost, no rate limits from an outside provider. The tradeoff is compute weight for a classification task that a smaller model might handle. The payoff is zero keys at risk, deterministic testability, and a classifier that can be interrogated locally before being trusted with live alerts.
The small-sample caveat. The current sample of vetted IDS verdicts is small. That is enough to validate the pipeline path, not enough to characterize false-positive or false-negative rates — both of which are labeled UNMEASURED. The alert path itself was verified end-to-end with a synthetic injection on 2026-07-13: a real event travelling from the IDS through the local classifier to push delivery, confirmed at the receiving endpoint.
Self-Healing Patterns
Three patterns compose to keep the system alive without human involvement. None are novel in isolation. What matters is how tightly they are specified and how faithfully they are documented when they fail.
Pattern 01 — Watchdog FSMs for daemon liveness
Each daemon runs as a supervised service with a watchdog that classifies its state into a finite set of transitions: idle → active → frozen → restarted. The FSM prevents restart-loops on transient states — "activating" and "heads-down conductor" are distinct from "frozen" and do not trigger recovery.
A critical production lesson: a liveness detector keyed on CPU and I/O counters reads GREEN on a reply-side wedge, because idle polling advances those counters even when the agent loop is frozen. The FSM distinguishes reply-delivery liveness from process liveness. Both signals are required; neither is sufficient alone.
Pattern 02 — Session-state resume after restarts
Each daemon owns a per-lane state file recording what it is executing right now and the next planned step. The file is written at step boundaries and cleared on completion — not every turn. On restart, the boot brief surfaces the state file first, so a restarted daemon resumes warm rather than cold-starting and dropping in-flight work.
The failure mode this closes: a daemon that hits a session limit mid-turn never auto-resumes after the window resets. Without the state file, queued work rots silently while every watchdog reads green — the worst failure mode because it has no alarm. With it, restart equals resume.
Pattern 03 — Guaranteed-delivery alert backstop
The alert path has a backstop that verifies delivery at the destination endpoint, not that the sending service is merely active. That distinction matters: a service can report "active," dispatch a payload, and still silently mis-route the alert to the wrong host or stream.
Verification method: synthetic end-to-end injection (performed 2026-07-13). A synthetic event traverses the entire pipeline — IDS tap through LLM classifier through push delivery — and is confirmed at the receiving device, not inferred from server-side logs. "Service active" is not the same as "alerts delivered."
Liveness is not "process alive." It is "the right output reached the right destination in the last N seconds." Every other definition is an optimistic proxy.
Incident Postmortems
Three real incidents from the permanent archive. Each produced an entry in an append-only lessons corpus — 160 entries spanning March–July 2026 — so the same failure cannot recur without a prior warning. The format is the differentiator: every incident resolves to a headline rule that the next session inherits before acting.
Stack
Result
What "healthy" means on this fleet is measured, not asserted. The numbers below come from the verified corpus — not from a monitoring dashboard that reports uptime without defining what it is counting.
The UNMEASURED column is not a failure — it is an honest instrumentation backlog. A metric labeled UNMEASURED tells you exactly where to invest next. A metric invented to fill the gap tells you nothing and erodes trust in everything beside it.
The archive holds 160 lessons in the postmortem format shown above, append-only, indexed by ID with a one-line headline rule. The index is regenerated on a schedule. Every correction becomes a rule the next session inherits before acting. Institutional memory does not expire when a model does.
Radical honesty beats managed optics. Missing data is more useful than invented data.