THE TOKEN GOVERNOR

Budget lanes, a reset-anchored spend-week, five downshift bands — the economics layer that keeps an agent fleet inside a fixed subscription

Systems ArchitectureSession Transcripts / Budget Lanes / TEMPO2025 — NOW

The Problem

An always-on agent fleet running against a fixed subscription has one failure mode nobody plans for: burning the week's budget in a night. Not a budget in dollars — a budget in compute allocation. Frontier models are heavy. Orchestration layers spawn sub-agents. Each planning pass, each verification loop, each independent check: tokens out. Nobody meters agents by default. They spend until the resource is gone, then stop, which is the worst possible time to stop.

The naive fix — rate-limit everything — just means slower work. The actual problem is the wrong model doing the wrong job. A frontier model writing a commit message. A frontier model formatting a JSON file. A frontier model doing the work a local open-weight model could do at zero marginal cost. Every unnecessary frontier call is borrowing from the judgment work that actually needs it.

Running a fleet this way is like a startup spending runway on catered lunches: not obviously wrong in isolation, catastrophic at the margin.

Token Governor system diagram showing four budget lanes feeding into five TEMPO downshift bands with a live fuel gauge overlay
SYSTEM DIAGRAM — BUDGET LANES AND DOWNSHIFT BANDS

What Exists

The governor is not a single script — it's an economics layer the fleet runs inside. Four interacting components:

BUDGET LANES

Four weighted lanes — operations, curation, builds, reserve — each with a unit allocation. A task declares its lane before it spawns. The lane determines its model ceiling, not the task itself.

SPEND-WEEK

The week is anchored to the subscription's actual reset day, not Monday. Heavy judgment work front-loads the fresh window. Mechanical chains ride the tail. The calendar week is irrelevant.

TEMPO BANDS

Five tiers keyed to percent of weekly pool consumed, intersected with a live 5-hour window band. Both dimensions must agree. A lean weekly band overrides a fresh 5-hour window.

FUEL GAUGE

A live readout every agent can query before spawning work. Weekly consumption %, 5-hour window %, and current effective band. The gauge is the contract: agents must check it, not assume.

Model weights underpin all of it. The frontier tier costs roughly 3.6× the baseline. The mid tier costs roughly 0.3×. A local open-weight model costs 0 marginal tokens — just electricity, which goes untracked (more on that in the honesty section). The governor uses these weights to convert band decisions into concrete model assignments, not suggestions.

Craft: The Downshift Cascade

The most important design decision in the governor is what happens when an agent asks for a model tier the current band doesn't allow. The naive answer: return an error. The correct answer: downshift silently and tell the agent what it got.

An agent asking for the frontier tier in a lean band gets the mid tier and a note explaining the band. It can proceed. If it genuinely needs frontier for a judgment call, it flags the exception explicitly — the governor has a judgment-exception path that bypasses the band check for specific task classes. The exception is logged, not silently granted. A cap self-declared by the spender is not a cap.

Frontier tokens buy thinking, cheap tokens buy typing. A cap self-declared by the spender is not a cap.

The five TEMPO bands, in order of spend pressure:

B01
FULL-PREMIUM< 40% of weekly pool consumed
Any model tier available. Frontier for judgment calls, mid tier for bulk.
B02
JUDGMENT-PREMIUM40–70% consumed
Frontier only for true judgment work. Mid tier becomes default for everything else.
B03
MID-ONLY70–85% consumed
Mid tier across the board. Frontier gated behind explicit judgment-exception flag.
B04
EMERGENCY85–95% consumed
Mid tier with hard token caps per task. No multi-agent spawning. Single-file writes only.
B05
LOCAL-TAIL> 95% consumed
Local open-weight model only. Zero marginal cost. Maintenance tasks only.

The 5-hour window band runs orthogonally: full-speed, wind-down, local-tail. The effective band is the intersection — the more restrictive of the two wins. At 68% weekly consumed and 9% of the 5-hour window used, the fleet runs in judgment-premium: frontier available but not default, mid tier as the workhorse. That was the live snapshot when this page was being built.

The Ledger

The governor is fed by a usage ledger parsed from on-box session transcripts: 48,769 rows of recorded usage, attributed by agent role and model tier.

48,769

usage rows parsed

8.7M

output tokens — top frontier consumer

8.1M

output tokens — second frontier consumer

6.5M

output tokens — third frontier tier

3.5M

output tokens — mid tier

11.8M

output tokens — planner role to date

Per-agent attribution is one of the ledger's design requirements: it must name the role that burned the tokens, not just the model tier. The planner role — responsible for decomposition, verification, and adversarial checks — accounts for 11.8M output tokens. That's the cost of thinking. The ledger names its own author.

The live snapshot at fact-gathering time: 5-hour window at 9% consumed, weekly window at 68% consumed. Effective band: judgment-premium. The site's case study pages were being written in mid-band, with the mid tier as the default builder model and the frontier tier reserved for verification and judgment calls.

The ledger names its own author. The planner role burned 11.8M output tokens to date — measured, attributed, visible.

Honesty

The governor has blind spots. The ledger reads only on-box session transcripts — usage from other surfaces (a mobile app, a web interface, anything outside the local session files) is invisible. Coverage percentage: UNKNOWN. The system could be measuring 60% of actual consumption or 90%. It cannot tell.

Weekly capacity is calibrated at a LOW-CONFIDENCE level. The bands were set from observation — watching what ran out when, adjusting thresholds accordingly — not from vendor-provided reconciliation data or a documented capacity contract. The governor is well-tuned for the workload it has seen. It is not validated against ground truth.

Local open-weight model electricity cost is untracked. The "zero marginal cost" framing is accurate for subscription metering; it is not accurate for total operating cost. The electricity draw of inference hardware running overnight is real. It goes into UNMEASURED.

The low-confidence calibration is known and documented — the fleet constitution calls it out directly. A cap self-declared by the spender is not a cap: the governor is aware of its own approximation.

UNMEASURED
—Coverage of on-box transcript scan vs total actual usage (UNKNOWN %)
—Weekly capacity against vendor ground truth — LOW-CONFIDENCE calibration
—Local model electricity cost — not tracked
—Cost-per-shipped-outcome vs single-model baseline — no controlled comparison
—Band calibration accuracy under workload patterns not yet observed

Stack

Session TranscriptsBudget LanesTEMPO BandsModel WeightsFuel GaugeSubscription Metering

Result

The fleet runs continuously. It has not burned a weekly budget in a single night since the governor was in place. The downshift cascade means work continues during lean bands — at reduced capability, not at zero. Emergency mode produces maintenance work. Local-tail mode keeps monitoring alive. The floor is never a wall.

The spend-week anchor matters more than the band thresholds. Anchoring to the subscription's actual reset day means the fleet front-loads judgment work when the pool is fresh and defers bulk work to the tail. A calendar-week anchor would misalign the heavy window with the actual available budget by up to a week.

The ledger's per-agent attribution has compounding value: when a workload category starts consuming unexpectedly, it surfaces immediately. Not "the fleet used more tokens than expected" — but which role, which task class, which model tier. The difference between a budget problem and an optimization target.

Treating tokens like startup runway changes how you allocate them: lanes over limits, downshift over cutoff, attribution over aggregate.

Cross-link: THE CONDUCTOR PATTERN — how the fleet uses the governor's model weights to route frontier tokens to thinking and mid-tier tokens to typing. THE INFERENCE RIG — the local open-weight infrastructure behind the local-tail band's zero-marginal-cost floor.