THE DECEPTION NET

Honeytokens, integrity monitoring, and injection detection — defending the AI agents themselves

Security EngineerHoneytokens / LLM classifier / hypervisor containment2026 — NOW

The Problem

When your operations team is a fleet of autonomous AI agents, the attack surface shifts. A traditional perimeter model assumes the humans are the targets. In an agentic system, the agents themselves — their memory, their prompt pipelines, their access to credentials — become the attack surface.

Three threat classes define the frontier: prompt injection (external content smuggled into agent instructions to redirect behavior), memory poisoning (corrupting the persistent notes and knowledge bases agents read before acting), and credential exfiltration (inducing an agent to echo secrets through any outbound channel). These are the new lateral movement. Standard endpoint tooling does not see them.

The studio runs a multi-agent homelab: several autonomous agents sharing a self-hosted virtualized infrastructure, local GPU inference, a remote operator console, and a growing corpus of persistent memory files. None of this is cloud-managed. None of it has a vendor's security team watching it. The defenses are whatever we build.

When the ops team IS the AI fleet, defending the agents is not a feature — it is the prerequisite for everything else working.

The Solution

The design answer is layered deception plus layered guards — not a single wall, but a set of detectors, sensors, and hard limits distributed across the stack. Each layer catches a different class of failure. Together they implement an assume-breach posture: accept that a sophisticated attacker might get one layer, build so that no single layer's failure is catastrophic.

Six layers are live. They run on the same infrastructure as the agents they protect — no separate security stack, no external service, no data leaving the homelab. Every classification call goes to the studio's own local inference model.

Six Layers, Live

01
HoneytokensBait credentials placed where only an intruder following lateral-movement patterns would look. If they are ever touched, the net has caught something real.
02
Memory Integrity MonitoringSensitive memory is checked against known-good baselines. An unexpected change is treated as a poisoning attempt. The alarm has never fired on a real intrusion, but the drills confirm it would.
03
Ingest GuardDeterministic provenance rules, followed by a local-LLM classifier for untrusted content. Harness tool envelopes are DATA; arbitrary web content is suspect. Both classification stages run locally — zero cloud exposure.
04
Credential-Egress GuardA pre-send hook scans every outbound agent message for secret-shaped content. Secrets never leave the host they live on, regardless of how the request is framed.
05
Write-Rights VerificationAutomated verification that each agent in the fleet writes only to its assigned paths, with any out-of-lane write surfaced to the operator.
06
Hypervisor vNIC KillAssume-breach containment: a detection event can cut a compromised VM’s virtual network interface at the hypervisor layer, out-of-band from the VM itself. The compromised machine cannot interfere with its own isolation.

Craft Details

Lesson 01 — Tool Envelopes Are Data, Not Directives

The local-LLM injection classifier worked well on its first pass — until we started seeing it flag content it should never have flagged: URLs, query parameters, and the fetch tool's own status messages. A line like REDIRECT DETECTED: following to https://... was being classified as a directive injected into the agent. It was, in the classifier's reading, an instruction embedded mid-content.

The fix required teaching the classifier that tool envelopes — the harness machinery that wraps every inbound observation — are part of the infrastructure, not external content. They arrive through a known structural channel with a known format. Classifying them as untrusted content is equivalent to flagging your own kernel as malware.

After the fix, we did not just re-run the happy path. We re-probed with a real injection payload — confirmed the classifier still caught it. A fix that tunes down sensitivity without a live adversarial re-check is not a fix; it is a quieter failure mode.

Lesson 02 — Trust the Class, Not the File

First-party security documentation — threat models, postmortems, hardening notes — is written to describe attacks. A naive content scanner reads those descriptions as attacks. The early ingest guard was flagging the studio's own security corpus as a prompt-injection corpus, because the text contained injection examples.

A per-file allowlist was the first instinct. It was the wrong one. A per-file allowlist rots: every new document in the corpus requires a manual update, and the guard fails open silently when an update is missed. One forgotten entry means the guard flags your own postmortem as hostile.

The correct fix is trust-classing by provenance. First-party documents are classed by where they come from and how they were authored, not by content and not by filename. Adding a new postmortem requires no allowlist update — it inherits its trust class from its provenance.

A per-file allowlist rots. A provenance trust class holds. The difference is whether the rule is about content or about provenance.

Design Note — Why Out-of-Band Containment

Every containment mechanism that runs inside the VM is, in a compromise scenario, controllable by the attacker. A compromised VM can suppress its own logs, kill its own monitors, and modify its own firewall rules. In-VM containment is better than nothing; it is not a hard boundary.

The hypervisor vNIC kill operates at a layer the compromised VM cannot touch. The hypervisor host can cut the VM's virtual network adapter regardless of what software is running inside. This is the containment that matters in an actual breach — and it is exactly the kind of control you only have when you own the hardware.

Defense in Depth

Stack

HoneytokensLLM ClassifierKVM HypervisorPython / systemdLocal InferencePre-Action Hooks

Result

All six layers are live. Alerts push to the operator's phone via a guaranteed-delivery path verified end-to-end with a synthetic event — we know the alert arrives on the real device, not just that the service is running.

The memory-integrity alarm has never fired on a real intrusion. That is the honest result: drill-tested, not battle-tested. The drills confirm the alarm would fire — we have not observed a live poisoning attempt. Attacker dwell time and formal false-positive / false-negative rates are UNMEASURED; the system is too new and the attacker sample size is zero.

What the deception net has changed is the asymmetry. Before these layers, a compromised agent could exfiltrate credentials, corrupt memory, and inject directives into sibling agents — and none of that would generate a signal. Now every one of those moves trips a different wire. The attacker has to be right every time. The defenses only have to catch them once.

The integrity alarm has never fired on a real intrusion. That is not a gap in the system — it is the system working.

The larger lesson: deception-based defense is cheap to instrument and expensive to evade. A honeytoken costs nothing to place. It produces zero false positives — nothing legitimate ever touches it. When it fires, the signal is unambiguous. The cost asymmetry is entirely in the defender's favor.