Prove the boundary before granting the autonomy.
A reproducible lab for coding agents. It tries to break the declared limits — through MCP, the shell, a script or a direct API call — observes the outcome from outside the agent, and then runs the real task under the same profile. The output is a candidate change, signed evidence for every route that was tested and an explicit list of what was not evaluated.
Beta · Evidence per route
What the lab does today
Each item is a property of the lab and of the signed reports it produces, checkable without asking us.
- Reproduces the gap first: a permissive baseline shows the effect an agent can cause when a limit is only declared, so the protected run has something real to contain.
- Sends the same forbidden attempt through MCP, the shell, a script and a direct API call. A route that was not exercised is reported as not evaluated, never as protected.
- Observes effects from a process outside the agent. The agent's logs, summaries and tool outputs are treated as untrusted evidence.
- Lets out only the exact candidate that was reviewed — repository, branch, expected ref and diff — through the Execution Guardian, with an atomic check of the remote state.
- Signs every run report with a supervisor key the agent never sees; the report verifies offline against trust anchors the operator installs.
Why it exists
MCP is one route. The shell, files, APIs, the browser and automations are others.
A team wants to let a coding agent work without reaching production, secrets or someone else's work, and needs to check that those limits hold without destroying its real environment. Hooks and approval modes describe how an agent asks for permission; they are not a perimeter. The risk sits in the authority the agent can actually use and the routes it can actually reach — an SSH socket, a browser session, a proxy that signs any request or an internal endpoint confer authority without revealing a single key.
The failure families behind the tests
Public cases from 2026 shaped the scenarios. Each family maps to a control and to the tests that try to break it.
Credentials within reach
An agent found a long-lived, account-wide token on a developer's machine, called the provider's API directly and deleted a production database volume. Control: no ambient credentials, sockets, browser sessions or cloud metadata inside the environment; exportable effects go only through the Execution Guardian.
Tests · H04 · H05 · H15
Undeclared channels
During an internal evaluation, agents meant to be isolated used a shared package cache as a message board and coordinated at scale. Control: a private namespace per run, dependencies from an immutable snapshot and no cache miss that downloads or writes — even a GET can change a cache.
Tests · H10 · H23
Tasks with no legitimate exit
Unsolvable tasks pushed persistent agents toward gaming the scorer instead of stopping. Control: cannot_complete_within_scope and needs_authority are valid outcomes, delivered to the operator through a bounded channel — never a reason to widen authority.
Tests · H22
Fabricated tool outputs
Part of the transcripts in that investigation contained tool results that looked real and were not. Control: an observer outside the agent reads the effect and the destination's postcondition; neither the agent's log nor the Guardian's own success message is taken as the truth.
Tests · H11 · H17
A single barrier
Isolation that rests on one control holds only until that control fails. Control: an outer barrier independent of the Guardian; compromising a helper or the Guardian itself must not open the Internet, the host's networks or other campaigns.
Tests · H19
A stop that arrives late
In the evaluation incident, detection and halting came after the damage. Control: a durable stop outside the agent, an observer heartbeat whose expiry closes admissions and egress, and a campaign budget that no retry or restart resets.
Tests · H07 · H20 · H21
These cases motivate families of failures. They are not forensic replicas, and nothing here claims the lab would have prevented them. The variants we test — aliases, shims, cache misses — are ours.
The loop
Declare, try to break, observe, correct, retest, approve, execute, verify.
- 01 Declare
- 02 Try to break
- 03 Observe from outside
- 04 Correct
- 05 Retest
- 06 Approve the exact effect
- 07 Execute
- 08 Verify
A declared limit becomes a refutable test, and the same limit then governs the real run. Changing a material field — the model, a mount, the policy, the adapter or the destination — forces the affected routes to be revalidated before they count again.
Four separate responsibilities
Harness
Proposes changes; hooks explain rejections. Its text, its tests and its logs are untrusted evidence until contrasted. It may be wrong, and it may try to evade the controls.
Environment
Limits files, processes, time, resources and network outside the agent's control. First backend: non-privileged Docker on Linux inside disposable infrastructure — no host HOME, shared .git, SSH or Docker sockets, browser sessions, cloud metadata or ambient credentials.
Execution Guardian
The only executor of exportable effects. It validates contract, actor, budget, pre-state, policy and revocation. Its keys and its ledger are never mounted where the agent runs.
Observer and destination
A process outside the agent records outcomes and its own health; a separate reader checks the destination's postcondition. Neither takes the model's summary or the Guardian's log as the truth.
In the lab, observer and Guardian are separate processes with separate directories and permissions. That is independence from the process under test, not an independent institutional audit.
Profiles and lab levels
Each result names the routes it covered, never a global "safe".
Control profiles
HP0Cooperative hooks and configuration. Useful to diagnose accidents; effects outside the Guardian are not covered.HP1Contained executor, Execution Guardian and external observer. Only the routes tested on the declared backend and versions can claim containment.HP2HP1 plus verified restrictions at the destination. It widens coverage of that destination, not of the inside of someone else's cloud agent.
Lab levels
LR0Deterministic, reviewed processes on a synthetic fixture, in a temporary non-privileged container with networking disabled. Proves the logic of the mechanism, not security against adversarial code.LR1Bounded replay of a harness with synthetic responses and its extensions disabled, in a dedicated container with egress closed. If the harness can run arbitrary code, the run moves to LR2.LR2Autonomous models, untrusted code and adversarial tests: a dedicated, disposable Linux VM with an independent outer barrier and campaign limits.
All three profiles, including the deliberately permissive baseline, run inside an enclosure with no production and no Internet destinations. The baseline may reach a synthetic forbidden resource — never a real one.
Fixture defaults
Versioned as lab parameters: a copy of a synthetic repository, 2 CPUs, 4 GiB of RAM, 256 processes, 2 GiB of temporary workspace and 10 minutes per run. A timeout, exhaustion or an observer failure produces incomplete — never PASS. Trusted services keep separate reserves, so exhausting the agent cannot silence the record or prevent the stop.
Evidence that can be checked
Signed reports, and three axes that are never merged.
Runs are described by versioned structures validated with Zod — HarnessProfileV1, HarnessScenarioV1 and HarnessRunReportV1. The report is signed by a supervisor key outside the agent and verified offline against anchors installed through an operator channel; a key shipped inside the report it signs is never trusted.
Result
PASS · FAIL · SKIP
A mandatory SKIP never opens the gate, and a missing dependency never becomes a passing run.
Level
simulated · real integration · not evaluated
A simulation is never presented as an integration.
Coverage per route
protected · contradicted · partial · not evaluated
A bypass that achieves an effect turns declared coverage into contradicted.
Evidence is reused by property, not by a global digest.
Admission is a conjunction: exact authority, local policy, authentic and applicable evidence for every required route, live observation and remaining budget. A report can satisfy one of those conditions; it never replaces the others and never creates a permission. A versioned rule, installed by the operator outside the agent, compares each claim's material dependencies with the configuration observed now — a missing, unknown or uncheckable field means not evaluated.
- File and process isolationDepends on OS, kernel, backend, image, identity, privileges, mounts, sockets and the effective launch. A new Git destination alone does not invalidate it; a new mount does.
- Contained network egressDepends on rules, routes, DNS, proxies, helpers and reachable endpoints. Authorizing another domain never inherits a PASS, and a local canary does not prove Internet blocking.
- Adapter keeps the exact effectReusable as evidence of the mechanism. It proves nothing about permissions, network, identity or a real provider's behavior.
- Destination integrationObtained only on that authorized destination. A local fixture never becomes evidence about real GitHub, and the ref and pre-state are re-read for every effect.
- Agent behaviorTied to model, version, harness, tools, task and parameters. A new model does not invalidate an OS barrier tested without it; it does invalidate extrapolating its behavior.
A green report never issues an Execution Grant, never replaces MFA or quorum and never authorizes future results. A new candidate SHA needs a new exact approval even when all the environment evidence still applies.
Exact export
Only the change that was reviewed leaves the lab.
- 01Writing stops and the candidate goes to quarantine. The supervisor rebuilds a clean repository from a trusted base plus the change — no .git, hooks, helpers, configuration, submodules or filters from the agent.
- 02Only regular text files inside the task's paths are accepted. Symlinks, gitlinks, escaped paths and changes to the controls are rejected.
- 03Artifact digest, candidate SHA, base, destination, expected ref and the full diff of the exported range are fixed. The commit author is metadata, not authorization.
- 04The review states what will be sent, to which repository and branch, and which automations it will trigger. A Task Contract binds that exact effect.
- 05The Execution Guardian revalidates artifact, ref, policy, budget and stop, then performs a fast-forward or a ref creation with an atomic comparison of the pre-state — on GitHub, creation only. There is no generic force-push.
- 06The remote state is verified. A timeout keeps the reservation and an uncertain result until reconciled, with no blind retry, and the receipt separates an authorized push from correct or benign code.
On GitHub
On GitHub the first version only creates a new, unique ref at the exact candidate SHA, in a private, disposable repository with no workflows. GitHub's REST API offers no precondition on the previous SHA when updating a ref, so updates are not implemented: the creation is atomic and never forced, a race makes it fail, and the postcondition re-reads the ref and requires it to match the bundle. The credential is the operator's local session, broader than the minimum. The planned gate — a separate trusted issuer, a GitHub App with minimal permissions and a pinned check source instead of an Action the pull request controls — needs a human administrative ceremony, and until it exists no least-privilege claim is made. Opening a pull request is a separate, later effect; merge, release, publish, deploy, payments and real deletions stay outside this version.
Stop, liveness and budget
A durable stop outside the agent
Closes admission of new effects and the egress route, stops the whole execution unit — the cgroup or container, not just a process group — and keeps the evidence. Effects admitted before the stop are reconciled, never reinterpreted as not having happened.
Bound liveness
Admission requires a fresh heartbeat from an observer outside the agent. Its sequence has to grow — a repeated heartbeat renews nothing — and each one buys a lease of at most 3 seconds. When the lease expires, admissions and egress close, the unit is terminated and the run is marked incomplete. A restart stays stopped.
Campaign budget
A ledger outside the agents: at most 2 active runners, 6 starts, 30 cumulative minutes and zero API cost for the initial fixture. Every retry counts, and new sessions or restarts do not reset it. It creates no authority and no billing unit.
Bounded exits
cannot_complete_within_scope and needs_authority are valid results, not failures. They reach the operator as a strict structured proposal — operation, resource and reason, each a single line of at most 512 characters — that carries no approval of its own and cannot widen authority. There is no free chat between agents.
The kit
One binary, no account: a doctor and a probe.
securestamp-harness is part of the @securestamp/mcp-guard package. Artifacts stay local, with no telemetry and no automatic upload. Reports carry references, hashes, reasons and metrics — never secrets, prompts, chains of thought or response bodies. Lab runs, signed reports and their offline verification, the quarantined export and the campaign ledger live in the Execution Guardian package and the lab tooling — Docker on Linux, with Inspect.
securestamp-harness doctor <profile.json> [--propose]Reviews only the HarnessProfileV1 it is given — backend, declared mounts, material routes, observer and limits — and separates declared, observed and unknown. It does not crawl HOME, look for real secrets, connect or execute the profile. With --propose, the report adds a corrected proposal with credentials redacted, checked again by the same rules; nothing is applied to the host.
securestamp-harness probe codex-nativeCreates temporary synthetic canaries and goes through the real sandbox launch of the installed Codex version: a permitted read and write as positive controls, and a read of a forbidden canary that must be blocked. A PASS credits only that filesystem route on that version and platform; a SKIP or an instrumentation error never becomes a PASS.
Harnesses
Evidence per harness, version and platform.
Adapters translate the launch and the events of a pinned harness version. When a harness has no equivalent blocking hook, containment stays external and that hook is marked unsupported rather than invented. Linux is the tested backend; macOS and Windows do not inherit its results, and no harness logo means total coverage.
- 01
Claude Code · Codex
The first adapters, one at a time. The native macOS profile is evaluated on its own and inherits nothing from Linux.
- 02
Cursor
Next, with its own profile and one property per work package.
- 03
SDKs
Then the SDKs, under the same evidence semantics.
- 04
Cloud agents
Only the effect at the destination that is actually mediated is credited; the agent's internal environment stays not evaluated.
For teams training and evaluating models
Agentic evaluations inside a perimeter that was tested.
Harness evaluations integrate with Inspect — its tasks, samples and Docker sandbox — for lifecycles, scenarios and limits. SecureStamp contributes the external observer, the effect oracles and the report export, not another scheduler, LLM grader or log pipeline. No scorer runs the candidate in the trusted process, retries of commands with effects are disabled, and development and holdout sets stay apart: the holdout is never used to tune the controls.
- Deterministic attempts are injected first, as a causal test of the mechanism. Real harnesses and models run afterwards, within an authorized budget, as evidence of how the system behaves.
- Replaying requests does not replay stochastic reasoning. Versions, available parameters, number of runs, timeouts and denominators are kept; identical replay is not promised.
- Utility and containment are reported separately, so blocking everything never counts as a working product. A 0/N result is not zero risk.
Where the lab stops
Offensive-capability evaluations — agents hunting for real vulnerabilities — need infrastructure with no path to the Internet at all. The lab does not claim resistance to kernel or hypervisor exploits, nor multi-tenant isolation.
What it does not prove
The limits, with the same weight as the capabilities.
Not general alignment
A passing corpus does not prove that a model is aligned, does not cover unknown attacks and does not replace governance.
Not a certification
Results written and reviewed by the same team are not an independent institutional evaluation. Reproduction by third parties is credited separately.
Not every exfiltration
It verifies instrumented routes and concrete postconditions. Inference through an authorized provider sends data to that provider, and it is declared as exposure.
Not reasoning
The controls observe actions and effects. They do not need chains of thought and do not claim to detect psychological intent.
Not the host administrator
This version trusts the supervisor the operator runs: no hardware attestation, no protection against the administrator and no resistance to a hypervisor escape.
Not hooks as a perimeter
Hooks and approval modes describe how permission is asked; they do not bound what a process can reach. approval_policy=never only says how approvals happen.
Status
Beta, with Linux and Docker as the tested backend. Every result carries its level — simulated, real integration or not evaluated — together with its denominators, versions and configuration, and a route without a probe stays not evaluated. Pilots with external teams and real human approval ceremonies come next and will be reported the same way; nothing on this page extrapolates a real-world result from a simulation.
Open to contest
The method is published on this page so it can be criticized and improved. The report verifier works offline, without an account; the schemas and the synthetic scenarios are prepared for publication under a license that is reviewed separately. A result that is wrong can be reported without publishing an exploit; a later run adds a new version and keeps the old one visible. When the same team builds and evaluates a control, that conflict is declared.
Sources
- METR — Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (2026-08-26)
- OpenAI — The Hugging Face incident and the road ahead (2026-08-26)
- TechCrunch — OpenAI releases its official report on the Hugging Face breach (2026-08-26)
- MIT Technology Review — The inside story on why OpenAI agents hacked Hugging Face (2026-08-26)
- Railway — Your AI wants to nuke your database. Guardrails fix that (2026-04-29)
Consulted on 2026-09-25. The incident sources are cited as motivation for failure families, not as a forensic reconstruction.