THE STREAMING HUNT

3 rounds. 3 falsified hypotheses. 3 real root causes.

Systems DebuggingLinux / GPU / Bootloader2024

The Setup

The rig: a Linux GPU workstation as the host, streaming its desktop at 5120×1440 at 120 Hz over a local network to a lightweight streaming client running integrated graphics. The goal was low-latency game streaming — the GPU lives in one room, the display and controls in another. The baseline was already elite: 2.9 ms round-trip latency. The streaming stack was mature. Nothing should have been wrong.

Three things were wrong. Each one had an obvious suspect. Each obvious suspect was innocent. This is the forensic record of finding the real causes.

Three symptoms. Three falsified hypotheses. The obvious suspect is rarely the cause.

The Three Rounds

ROUND 01

THE FPS COUNTER THAT LIED

SYMPTOMDesktop showed 68–72 fps. Looked like a cap. Host GPU: 0–3% render, idle.
OBVIOUS SUSPECTA frame-rate cap or throttle somewhere in the streaming stack.
FALSIFICATIONA constant-motion test page — browser window full of spinning content — jumped instantly to 120. No cap. No throttle.
ROOT CAUSEThe streaming stack only transmits frames when pixels change. An idle desktop at 70 fps was the system working correctly: 70 on-screen motion events per second, not 70 out of a possible 120. The counter measured motion, not capability.
SUSPECT FALSIFIED
ROUND 02

THE GREEN SCREEN THAT WASN'T THE ENCODER

SYMPTOMIntermittent green flashes under load — full-frame corruption, unpredictable.
OBVIOUS SUSPECTThe host encoder. Tuning encoder settings (bitrate, codec params) made things worse.
FALSIFICATIONDeeper encoder tuning produced no improvement — it increased corruption frequency. The server-side encoder was not the variable.
ROOT CAUSEThe client's own kernel log held the answer: five GPU HANG resets in a ~400-second session. The integrated hardware decoder on the streaming client was hitting its stability limit — roughly 885 megapixels per second at 5120×1440×120 exceeded what that kernel's firmware could sustain. A newer kernel enables the hardware-decode stability layer. After reboot: zero hangs, zero green frames.
SUSPECT FALSIFIED
ROUND 03

THE KERNEL THAT WAS INSTALLED BUT NEVER BOOTED

SYMPTOMThe fix was applied — newer kernel installed — but the machine kept booting the previous version.
OBVIOUS SUSPECTInstallation failure. Package conflict. Some kernel module didn't land.
FALSIFICATIONPackage manager reported a clean install. No errors. The kernel was on disk.
ROOT CAUSEA leftover bootloader pin: GRUB_DEFAULT hard-pointed at the old kernel entry by an earlier unfinished configuration session. The boot menu was hidden and timeout was set to zero — no UI to catch it, no pause to notice. The new kernel existed but was never offered. Repointed the default, restored the menu with a short timeout, took a backup of the config.
SUSPECT FALSIFIED

The Discipline Detail

The baseline before any of this started: 2.9 ms round-trip latency. That is, by most measures, already elite for local game streaming. During the hunt, several theoretical “optimizations” were applied — encoder parameter tuning, network buffer adjustments — measured carefully, and found to perform worse: 4 ms. They were reverted the same session.

The post-fix result, with the client's decoder stabilized and the correct kernel actually booting: sub-2 ms sustained. The improvement came from fixing the real problem, not from chasing an already-elite metric.

A separate finding during Round 2: the host encoder was running at 98% saturation, outputting approximately 130 Mbps. This wasn't the cause of the green flashes — but it was a real ceiling worth backing off. Found during the hunt, resolved as a separate adjustment.

Don't chase an already-elite metric. Find the actual problem. The improvement arrives when the cause is fixed, not when the number is tuned.

The Forensic Discipline

The methodology held across all three rounds: collect concrete evidence first, form one specific hypothesis, design the smallest possible test to falsify it, run the test, accept the result. When the test falsifies the hypothesis — and it did all three times — return to evidence, not to intuition.

What Generalizes

Measure the real quantity. A metric label is not a guarantee that it measures what you think. The fps counter measured on-screen motion — correct for its design, wrong for what the diagnosis needed. Before trusting a diagnostic signal, ask: what exactly does this measure? For what definition of the term it's named after?

Attribute corruption to the consumer, not the producer. Intermittent pixel corruption under load is more likely to be a decode instability at the receiving end than an encode problem at the sending end. The server was running fine. The client was dropping frames in its decoder. Look at where the artifact appears, not where the data was generated.

Installed ≠ booted. A package manager confirming a successful install means the files are on disk. It says nothing about the boot configuration that decides which kernel actually loads. When a fix doesn't take effect after a reboot, read the bootloader config before declaring the install broken.

Installed does not mean booted. Encoded does not mean decoded. A metric named for a thing does not always measure that thing.

Stack

Linux / Kernel DebugGPU Decode DiagnosticsGRUB BootloaderGame Streaming StackGPU Workstation Host

Result

MEASURED
Debugging rounds3
Hypotheses tested3 (all falsified before fix)
Wall time~1 hour
GPU HANG events — before fix5 in ~400 seconds
GPU HANG events — after fix0
Latency — baseline2.9 ms
Latency — bad experiment peak4 ms (reverted same session)
Latency — post-fix sustainedsub-2 ms
Host encoder saturation found~98% / ~130 Mbps (adjusted)
UNMEASURED
Recurrence sampleOne client, one fix — no longitudinal data
In-game frame consistencyDesktop streaming verified; in-game not separately measured

The rig has been stable since. The streaming stack that was exhibiting three separate failure modes is now operating at sub-2 ms latency with zero decoder hangs. The fixes were targeted: one kernel update, one bootloader repoint, one encoder parameter pullback. Nothing about the host GPU, the network, or the streaming software itself required modification.

The discipline — evidence before hypothesis, one hypothesis at a time, smallest possible falsifying test — is what kept the investigation from sprawling into weeks of misattributed tuning. Each round took less than twenty minutes once the right question was isolated.

Three rounds, ~1 hour, zero recurrence. The constraint wasn't compute — it was asking the right question before writing the first fix.

Related Work

The same forensic discipline — evidence before hypothesis, device evidence over emulator evidence — drove the cross-browser debugging story in The Hero Rebuild. That investigation ran seven rounds across three browsers before a single screenshot from the real device ended it.

The self-healing infrastructure that catches system failures automatically is documented in The Self-Healing Lab.