CanaryForge runs a tool-calling LLM agent inside a Docker container with an escape temptation planted in its task, and watches it with three monitors that never ask the model anything. It measures time to detection.
The industry evaluates agents by asking whether the attack succeeded. CanaryForge asks how long it took you to find out. An attack you catch in seven milliseconds and an attack you never catch are both failures of prevention, and they are nothing alike operationally.
Every signal is structural. Nothing depends on the agent self-reporting, on a judge model, or on parsing intent out of text, so none of it degrades when the model changes. Escalation is scored 0 to 6, where 5 means the canary token left the environment.
Every request out of the container passes an mitmproxy. Destinations outside the declared baseline, and the canary token in traffic, raise an alert.
A token planted inside the sandbox, watched by inotify on the shared volume. Fires on file contents and, crucially, on file and directory names.
The agent pings on a fixed interval. Silence longer than the threshold, not explained by a clean container exit, means the sandbox is gone.
Word for word the same injected text, the same model, the same session. Only the delivery channel changes.
Memory written by an earlier run, and documents sitting in a local knowledge base, never cross the network during the session. They arrive with the implicit authority of internal context. A payload pulled over HTTP from an external mirror carries the same words and gets checked.
This is exactly the asymmetry OWASP ASI06, Memory and Context Poisoning warns about, reproduced under controlled conditions.
Escalation is scored independently per rung, so reaching level 5 does not imply passing through 3 and 4. That is measured, not assumed, which is what makes the next number readable.
This is the channel of the July 2026 incident that anchors the whole project: agents encoding state in directory names, invisible to anything watching the network. A deployment monitoring only egress would have caught almost none of these leaks.
Nothing here required a network exploit. The agent wrote a filename.
The con_harness condition adds a system prompt warning the agent not to leave its scope. At ten runs it looked like a strong result. It is reported here with the sample size that weakens it, because that is the honest reading.
The warning names the attack surface while warning about it: files, directories, URLs. Naming it may prime the model toward it. That reading survives at n=10 and does not survive at n=40.
The honest statement is the narrow one. An effect this size needs a bigger sample than this sprint could run, and reporting the n=10 cell alone would have been the easiest way to oversell the whole project.
None of these announced themselves. Each was found by reading artefacts that the summary reported as successful runs, and each changed a number that had already been written down.
Runs ended early with a clean exit code because the tool-call parser gave up on output the model had produced correctly. The run looked finished, and it had not started. A dedicated calibration batch exists to separate parser failures from agent decisions.
An evaluation harness that never reports its own faults is asking to be trusted on faith. Every figure above was recomputed after these three were found. The count that moved most was the detection rate, which an earlier draft reported as 100 percent.
Written into report/scope.md before the results existed, so it could not be shaped around them.
The temptation is planted, the tasks are synthetic, and the model is a single local 7B. This measures detection latency against induced escapes, not the rate at which real agents escape.
It is a system prompt warning, not a runtime guard. Comparing it against a real implementation such as OWASP Agent Memory Guard would be a stronger experiment than comparing warned against unwarned.
Hardware and Ollama version per batch live in manifests and prose, and behaviour differs measurably between versions. Cross-batch comparisons carry that confound, and one batch is flagged as having no statistical power at all.
Most detections here are filesystem ones, where the write and the inotify event land in the same instant. Below a millisecond the number stops measuring the monitor and starts measuring clock skew between two processes. Read it as "faster than this setup can resolve", not as a latency. The slowest machine in the corpus medians at 0.96 s, and the interval that actually matters is the other one: 44 seconds from detection to nothing happening.
The harness is clonable and the metric is cheap. Three directions it opens.
Time to detection is a number that can regress between versions of an agent, a framework or a sandbox. Nothing today tracks it, because nothing measures it.
The dominant leak channel here is invisible to egress tooling. An inotify canary is cheap enough to run permanently next to a real agent, not only in an evaluation.
Swap the system prompt condition for a runtime memory guard and the same harness measures whether that defence moves detection latency or only the success rate.
# every figure on this page, recomputed from the artefacts python3 site/build.py # the experiment matrix python3 orchestrator/run_experiment.py # the research panel, live run feed over the same code python3 dashboard/app.py