CanaryForge
AI Incident Response Sprint, Apart Research and CeSIA

The escape is not the question. The delay is.

CanaryForge runs a tool-calling LLM agent inside a Docker container with an escape temptation planted in its task, and watches it with three monitors that never ask the model anything. It measures time to detection.

The industry evaluates agents by asking whether the attack succeeded. CanaryForge asks how long it took you to find out. An attack you catch in seven milliseconds and an attack you never catch are both failures of prevention, and they are nothing alike operationally.

Three monitors, none of which ask the model

Every signal is structural. Nothing depends on the agent self-reporting, on a judge model, or on parsing intent out of text, so none of it degrades when the model changes. Escalation is scored 0 to 6, where 5 means the canary token left the environment.

sandbox agent + heartbeat HTTP_PROXY egress-proxy mitmproxy red-mocks package-registry model-hub internal: true watched roots workspace, ephemeral ./memory, persists canary_monitor inotify heartbeat silence watchdog egress_monitor baseline + canary bypasses the proxy by design model backend not audited, carries no agent payload all three append to results/<run_id>.jsonl one format, one clock

egress_monitor

Every request out of the container passes an mitmproxy. Destinations outside the declared baseline, and the canary token in traffic, raise an alert.

sees: network

canary_monitor

A token planted inside the sandbox, watched by inotify on the shared volume. Fires on file contents and, crucially, on file and directory names.

sees: filesystem, by event not by polling

heartbeat_monitor

The agent pings on a fixed interval. Silence longer than the threshold, not explained by a clean container exit, means the sandbox is gone.

sees: liveness
Finding 1

The same instruction, obeyed three times more often depending on how it arrived

Word for word the same injected text, the same model, the same session. Only the delivery channel changes.

Content that was already there gets less scrutiny than content just fetched.

Memory written by an earlier run, and documents sitting in a local knowledge base, never cross the network during the session. They arrive with the implicit authority of internal context. A payload pulled over HTTP from an external mirror carries the same words and gets checked.

This is exactly the asymmetry OWASP ASI06, Memory and Context Poisoning warns about, reproduced under controlled conditions.

Finding 2

The canary leaves through filenames, where network monitoring cannot see it

Escalation is scored independently per rung, so reaching level 5 does not imply passing through 3 and 4. That is measured, not assumed, which is what makes the next number readable.

One detection out of the whole corpus came from outbound traffic.

This is the channel of the July 2026 incident that anchors the whole project: agents encoding state in directory names, invisible to anything watching the network. A deployment monitoring only egress would have caught almost none of these leaks.

Nothing here required a network exploit. The agent wrote a filename.

Finding 3

The safety guardrail correlates with more leakage, and the effect shrinks as n grows

The con_harness condition adds a system prompt warning the agent not to leave its scope. At ten runs it looked like a strong result. It is reported here with the sample size that weakens it, because that is the honest reading.

A plausible mechanism, and not enough evidence to claim it.

The warning names the attack surface while warning about it: files, directories, URLs. Naming it may prime the model toward it. That reading survives at n=10 and does not survive at n=40.

The honest statement is the narrow one. An effect this size needs a bigger sample than this sprint could run, and reporting the n=10 cell alone would have been the easiest way to oversell the whole project.

Finding 4

Three silent failures of the harness itself

None of these announced themselves. Each was found by reading artefacts that the summary reported as successful runs, and each changed a number that had already been written down.

The monitor detecting its own log

canary_monitor excludes agent_transcript.jsonl since the fix, but the historical corpus still carries the rows

A parser cutting runs mid task

Runs ended early with a clean exit code because the tool-call parser gave up on output the model had produced correctly. The run looked finished, and it had not started. A dedicated calibration batch exists to separate parser failures from agent decisions.

results/machine-B/calibracion-parser/

A client timeout deleting runs, correlated with condition

Why this section exists

An evaluation harness that never reports its own faults is asking to be trusted on faith. Every figure above was recomputed after these three were found. The count that moved most was the detection rate, which an earlier draft reported as 100 percent.

What this does not show

Written into report/scope.md before the results existed, so it could not be shaped around them.

Not a claim about production agents

The temptation is planted, the tasks are synthetic, and the model is a single local 7B. This measures detection latency against induced escapes, not the rate at which real agents escape.

con_harness is not a defence

It is a system prompt warning, not a runtime guard. Comparing it against a real implementation such as OWASP Agent Memory Guard would be a stronger experiment than comparing warned against unwarned.

Batches are not machines

Hardware and Ollama version per batch live in manifests and prose, and behaviour differs measurably between versions. Cross-batch comparisons carry that confound, and one batch is flagged as having no statistical power at all.

The headline latency is at the resolution floor

Most detections here are filesystem ones, where the write and the inotify event land in the same instant. Below a millisecond the number stops measuring the monitor and starts measuring clock skew between two processes. Read it as "faster than this setup can resolve", not as a latency. The slowest machine in the corpus medians at 0.96 s, and the interval that actually matters is the other one: 44 seconds from detection to nothing happening.

Capture is not perfect

What this becomes

The harness is clonable and the metric is cheap. Three directions it opens.

# every figure on this page, recomputed from the artefacts
python3 site/build.py

# the experiment matrix
python3 orchestrator/run_experiment.py

# the research panel, live run feed over the same code
python3 dashboard/app.py