THE HONESTY GATE
An AI strategist whose confidence is mathematically capped by its evidence — and graded afterward
The Problem
AI advice fails most often not through bad reasoning, but through fake confidence. A model with half the relevant data will still return a recommendation at 0.92 confidence if nothing stops it. The structural problem is not that models lie intentionally — it's that nothing in the generation path connects the strength of an assertion to the quality of the evidence beneath it.
A daily strategist that you can't trust to know what it doesn't know is worse than no strategist. It trains you to ignore caveats. It conflates missing data with confident irrelevance. It produces a 0.90-confidence recommendation on a day when the metrics digest was unreachable and three key backlog items were stale — and it presents that recommendation exactly like the one from the day everything was current.
The problem is tractable, but only if you build the ceiling into the pipeline, not into the prompt.
The Strategist
The system produces ONE recommendation per day: a single "do this today" action drawn from an 11-source briefing packet. Sources include the metrics digest, current backlog state, financial indicators, a research drop, activity logs, and related feeds. Any source that fails to load enters the packet honestly — as available: false with a value of "unknown". Nothing is fabricated to fill a gap.
The recommendation is structured into 10 required fields: action, why, what-it-unlocks, what-to-ignore, risks, the 30-minute version, the 2-hour version, done-definition, confidence, and evidence-used. It doesn't ship until it passes the gate.
Unknown stays unknown. A missing source enters the packet as absent — never as a best guess the model fills in.
The Gate
The gate is deterministic. It does not ask the model to evaluate its own honesty — it measures the recommendation against rules the model cannot override.
★ CORE RULE — COVERAGE CAP
max_confidence = 0.30 + 0.70 × (sources_available / sources_total)
Six sources of eleven? Your ceiling is 0.68 — no matter how sure you feel.
All eleven available? Ceiling is 1.0. That's the only way to get there.
The full gate runs five checks in sequence:
- Structure: all 10 required fields must be present and non-empty.
- Range: confidence must be a number in [0, 1].
- Coverage cap: confidence may not exceed
0.30 + 0.70 × (available / total). - Evidence citation: every source named in evidence-used must be a source that was actually available in this run's packet. No citing what wasn't there.
- High-confidence bar: confidence ≥ 0.5 additionally requires at least 2 distinct evidence sources. A high-confidence recommendation backed by a single data point is rejected.
★ GATE SELFTESTS — 9/9 GREEN
All 10 required fields present
Baseline structural pass
Missing required field (action omitted)
Field guard catches it
Confidence = 1.3 (out of [0,1])
Range guard catches it
Confidence = 0.95, 6 of 11 sources available
Cap = 0.30 + 0.70 × (6/11) = 0.68 — 0.95 exceeds ceiling
Confident-but-thin: confidence = 0.55, evidence_used cites unavailable source
Evidence citation guard: can only cite what arrived
Confidence = 0.5, only 1 evidence source listed
High-confidence requires ≥2 distinct sources
Humble low-confidence: 0.35 with 3 of 11 sources, 2 evidence entries
Cap = 0.49 — 0.35 is safely below; evidence count satisfies the low-confidence path
All 11 sources available, confidence = 0.70
Full coverage → ceiling is 1.0; 0.70 passes
All 11 sources available, confidence = 0.70, evidence cites fabricated source
Evidence must cite from the ACTUAL available set, not invented names
Craft: The Audition That Shaped It
Before the frontier tier took the strategist role, a local open-weight model was tested. It was fast, free to run, and produced structured JSON that parsed cleanly. The audition ran live against a real briefing packet.
The model returned a recommendation at 0.92 confidence. The coverage cap caught the first problem: only 5 of 11 sources were available that day, putting the ceiling at 0.65. Rejected on structural grounds.
But the structural rejection masked a deeper problem. Reading the rejected output, the model had fabricated product specifics and sales figures that appeared nowhere in the briefing packet — inventing plausible-sounding details to fill the gaps where sources were missing. The structural gate caught the confidence violation. It did not catch the invented prose.
A structural honesty gate does not catch in-prose fabrication. That became a permanent lesson. Synthesis stays on the frontier tier; the gate stays as the floor, not the ceiling.
This distinction matters: the gate rules out the provably wrong (overconfident on thin data, citing unavailable sources, missing required fields). It cannot rule out the subtly invented. That's a model-quality problem, not a validation problem — and the correct answer is to route synthesis to a model tier that doesn't hallucinate the gaps, while keeping the gate as the structural floor that never relaxes regardless of which model is upstream.
Craft: The Prediction Ledger
A strategist that never gets graded is a strategist that never gets better. Every recommendation that passes the gate is logged to a prediction ledger before the outcome is known: the expected score, the reasoning, and a confidence reading — all written at decision time, before the week plays out.
Graded outcomes update a belief store. Predictions that land accurately strengthen the weight of the reasoning patterns that produced them. Those that miss, weaken them. The calibration is meant to compound over time.
205
Predictions logged
18
Graded to date
1.5 pts
Avg error (10-pt scale)
8.8%
Grading coverage
What the Numbers Don't Say Yet
The page must say this plainly: the calibration data is thin, and what exists is biased.
8.8% grading coverage
18 of 205 predictions have outcomes logged. The calibration curve is accumulating. Conclusions drawn from it now have a denominator of 18, not 205 — and that denominator should accompany every verdict.
13 pre-negotiated entries
Of the 18 graded entries, 13 came from cases where the principal had already signaled intent before the recommendation was delivered. Those entries have a 100% hit rate — but that measures packaging quality, not predictive ability. They are not evidence of calibration. They are evidence that the strategist can recognize and articulate a decision already in progress.
Wiring gap found during this fact-gathering
The evidence-linkage field — which is meant to connect each ledger entry to the specific sources it cited — went unfilled for the first 17 runs. This was discovered during the research pass that produced this page. The field exists in the schema. The pipeline did not write to it. Those 17 entries cannot be retroactively audited for source-to-conclusion traceability. That's reported here rather than quietly omitted.
The house rules that govern this system — "a denominator accompanies every verdict" and "UNMEASURED beats invented" — apply to the system's own self-description. This page is written under the same gate it describes.
Stack
Result
The gate has been live since early June. Every recommendation that ships has passed five deterministic checks; recommendations that don't pass are logged with their rejection reason rather than silently dropped or retried without bounds. The local model was rejected after one audition session. Synthesis has stayed on the frontier tier since.
The prediction ledger is accumulating. The calibration signal will be meaningful when the grading coverage is a real denominator rather than 18 of 205. The 1.5-point average error on a 10-point scale is a real number from a small sample. It's reported with that caveat, not without it.
UNMEASURED
- —Decision quality improvement from the ledger grading loop (no controlled baseline)
- —False rejection rate — how often the gate rejects a genuinely good recommendation on a coverage technicality
- —Coverage curve over time — whether the 8.8% grading rate is increasing or plateaued
- —Belief-store delta effectiveness — whether strengthened/weakened patterns actually change recommendation quality
Fake certainty gets rejected at the gate. Real uncertainty ships with its denominator. That's the whole design.