THE CONDUCTOR PATTERN
One frontier conductor. Waves of mid-tier builders. Adversarial verification at every gate.
The Problem
Run one expensive model on everything and you burn your budget on boilerplate. It types fast, costs a lot, and doesn't get smarter because it has more money — it just gets more expensive. Worse, a single model doing everything is serial by nature. You wait.
Run an unsupervised swarm of cheap models and the output is slop at scale. Cheap models hallucinate metrics, invent dates, contradict each other, leak private details into text they shouldn't, and verify their own work with the same mind that wrote it. A swarm without oversight produces high throughput and low trust.
The real failure mode is subtler than either: a builder subagent reviewing its own output is not a verifier — it's a self-congratulation loop. The party whose work is judged cannot write the judge's instructions alone. You need independence at the gate, and independence is hard to buy cheaply.

The Pattern
The conductor pattern separates thinking from typing at the model tier. One frontier-tier agent — the conductor — does the work that requires judgment: gathering facts, writing specs, decomposing tasks into clean scoped units, dispatching builders with exact instructions, then running every verification gate itself after the builders return.
Mid-tier builders receive a spec, write one file each into clean context, and return. They never touch git. They never deploy. They never run the test suite. They never verify their own privacy compliance. The conductor does all of that — independently, after the fact, with a mandate to refute rather than confirm.
Frontier tokens buy thinking. Mid-tier tokens buy typing. They are not interchangeable — and treating them that way wastes both.
The conductor's first act on any build wave is to gather ALL facts before a single builder receives a prompt. Numbers are verified on disk, UNMEASURED labels are applied to anything unverified, and the fact packet is locked. Builders may not invent metrics. They receive one fact packet and one file assignment. If a number isn't in the packet, it doesn't appear on the page.
The conductor's last act on every wave is to run every gate personally: an independent privacy grep on each page (never delegated back to the builder who might miss their own leaks), a typecheck pass, a production build, dual-engine render verification in both WebKit and Chromium, screenshot review, then commit. A builder's output that fails any gate goes back for revision before it touches the repository.
Craft Details
The independence rule. A builder's own subagent is never its verifier. The verification mandate is REFUTE — stated explicitly in the conductor's prompt to every checker. A verifier whose mandate is “confirm this looks good” is not a verifier; it's a rubber stamp with extra compute cost. The conductor writes the verifier's instructions, not the builder. The builder never touches those instructions.
Clean-hands dispatch (A/B measured result). An earlier approach loaded builders with heavy scaffolding context — prior pages, design discussion, partial architectural history. It seemed helpful. It wasn't. A controlled A/B evaluation compared scaffold-heavy context against clean dispatch (spec only, one file assignment). Clean-handed dispatch won by approximately 7 points on a qualitative build-quality rubric. The scaffolding was noise. The builder performs better when it has exactly what it needs and nothing it doesn't.
The 5%/95% ratio as a quality rule. The conductor (frontier tier) consumes roughly 5% of total tokens per wave. Mid-tier builders consume the other 95%. This ratio is not a cost target — it's a quality signal. A frontier model reviewing ten cheap builds compounds: it catches the same class of mistake across every file without burning the full frontier rate on the typing. A frontier model typing one build doesn't compound; it just spends. The ratio tells you which role the model is playing. If your “conductor” is at 50% of tokens, it's a co-builder, not a conductor.
Writers write incremental artifacts. Every builder writes its output file to disk early — before it is complete — and improves in place. This is not style; it's insurance. A mid-flight API error on a builder that has only written to context loses everything. A builder that wrote to disk early loses at most the last few paragraphs. The conductor's dispatch prompt includes this requirement explicitly.
The Receipt
This site was built with the conductor pattern. Today. The numbers below are the actual build log from this afternoon — not an illustration of the pattern, not a hypothetical. The receipt.
Zero render errors across dual-engine verification (WebKit + Chromium). Zero privacy leaks across every page — every page received an independent privacy grep by the conductor before commit, not self-certified by the builder who wrote it.
Including this page. The page you are reading was written by a mid-tier builder from a conductor's verified fact packet. The conductor ran the independence check, the typecheck, and the render pass before it touched the repository. The builder never saw git. That's the pattern. This is its own receipt.
The best proof of a system is the artifact it produced while proving itself. This page is both.
Stack
Result
15 case-study pages live on this site as of today. Three build waves, approximately 110 minutes of wall time, zero render errors, zero privacy leaks. The conductor reviewed and committed every page; no builder touched the repository directly.
The A/B measurement confirmed what the pattern's logic predicted: clean-context dispatch outperforms scaffold-heavy context by approximately 7 points. The 5%/95% frontier-to-mid token ratio held across all three waves.
The independence rule held throughout: every verification gate was run by the conductor, not the builder whose output it was checking. No page was self-certified.
UNMEASURED
- Cost-per-shipped-page versus a single-model baseline — no controlled comparison run.
- Builder error rate beyond today's sample — 0 errors in 15 pages is a real result, but the N is small.
- Long-run quality drift as spec complexity grows — the A/B was one session, not a longitudinal study.