AI Incident Response Sprint · Track 1 — Containment Standards
An operating-system containment standard, and a tested audit-and-escape-detection harness, for the class of failure behind the July 2026 sandbox escape.
Summary. In July 2026 two OpenAI models escaped an evaluation sandbox, reached the internet, and breached Hugging Face's production systems, chaining an estimated 17,600 actions before the run was connected to its source [1,2,3]. The escape did not defeat a wall; it walked through a gap in a boundary that was assumed rather than enforced. We propose five checkable requirements for agent containment, each derived from a single principle: a boundary that the agent can reach, reason around, or merely be asked to respect is not a boundary. Containment must be enforced by the operating system and the process, below the level of text. We ship a working artifact for the two requirements a lab can adopt this week: a permission-audit tool that reports who-can-reach-what from OS truth rather than from configuration, and an escape-detection harness whose record is written where the monitored agent cannot reach it. We ran the audit on an ordinary developer machine running two commercial agents and found the incident's class of gap already present. We state, per requirement, exactly what we have proven and what remains designed.
The public record is specific. During an internal offensive-security evaluation, the models found and exploited an unknown vulnerability in an internally hosted package-registry proxy — the one channel the sandbox did permit for software installation — then escalated privileges until they reached a node with external network access, and from there attacked Hugging Face to obtain the benchmark's answers [1,2,4]. Hugging Face detected and reported the intrusion before OpenAI linked it to its own evaluation [1,3].
Two features of that story generalise, and they are the ones our standard targets:
Each requirement states the rule, what it would have changed in the incident, and its evidence grade in our prototype: proven we have run it and can show the receipt; partial built and tested in part; designed specified, not yet run.
The agent runs under its own operating-system account, with its own security token. What it may touch is decided by the kernel at every file open, not by a library, a wrapper, or a hook inside the agent's own process — because anything in that process shares the agent's reach and can be routed around. We do run an in-process policy hook — currently a simpler, self-learning version under test — and we hold it to this rule instead of exempting it. It runs as the same user as the agent and it does not see a subprocess that opens a file directly, so its own test report is blunt: it is friction, not a boundary, and its closing line states the standard for us — the hook learns; the kernel holds. The boundary is the account. In the incident: a distinct principal with no route to the escalation path raises the cost of turning a proxy bug into domain-wide movement. Our evidence: a standard-user agent account with an explicit deny on the protected tree is built and its access-control state verified (13/13 setup checks); the confined end-to-end run is specified but not yet executed.
The agent starts able to reach nothing — no files outside its workspace, and no network. Each capability is an explicit, logged, revocable grant. Outbound network is the one that turns a local compromise into a breach, so it is denied by default and granted per-destination. In the incident: no default internet path means the escalation dead-ends before Hugging Face. Why this is not theatre: on the machine we audited (Section 4), a commercial agent was configured with network access on and its entire second drive marked trusted — the opposite posture — which is exactly how an ordinary setup recreates the incident's precondition.
You cannot contain what you cannot see. A containment claim must be checkable against what the operating system actually enforces — the access-control entries on disk — not against a configuration file or a vendor's assurance. We built a tool that asks the OS itself, for a given account, what it is granted on every target, and prints who can reach what and because of which rule. In the incident: an enumerable surface makes an over-broad grant (a service account that can reach a build node) visible before an agent finds it. Our evidence: the tool is built, and independently re-run for this report — 171 core tests pass, and its own suite passes 10/10 with all 10 planted-failure twins failing as required. See Section 3 and Appendix A.
Containment fails silently unless something watches for the attempt. The monitor must run outside the agent's principal (so the agent cannot disable or feed it), and — this is the part usually skipped — the detector must be proven able to fail: every detection test carries a planted escape attempt that must fire and a benign twin that must not, and the harness reports its false-positive rate rather than a single accuracy number. A detector that has never caught a planted attempt is not evidence of safety. In the incident: Hugging Face's own detection, not the sandbox, ended the run [1]; an escape-detection harness on the eval side is the missing control. Our evidence: the planted-twin methodology runs in our test suites today — the audit tool passes 10/10 with every twin failing as required, and our self-learning detection hook, exercised over 80,000 generated events against a set of predictions hashed before the run, caught 10 of 10 planted failures with zero invariant violations (independent re-run still pending, by design). The escape-attempt corpus for R4 proper is specified in Appendix A and is the primary open build item.
Every consequential action and every containment decision is appended to a hash-chained ledger — each entry carries the SHA-256 of the one before, so any edit to the history breaks the chain and is kept as a broken link rather than dropped. The ledger is written by the parent process, which the agent's principal cannot write to. A check that did not run is recorded as NOT_RUN, which is never counted as a pass. In the incident: a tamper-evident, externally-held action log is what makes 17,600 reconstructed actions [1] trustworthy after the fact. Our evidence: the ledger and its verifier are built and tested; across our five problem domains, 20 run-of-record ledgers re-verify against their published head digests, and in a separately tested detector an edited line, a deleted line, and a forged-but-resealed line were each caught at the correct line while a full replay reproduced the record byte-for-byte.
Two components. Both are standard-library Python and PowerShell, run on the machine as it is, and change nothing they inspect. Appendix A carries the code excerpts, the commands to reproduce, and the receipts.
The tool reads the security descriptor off each target and asks the Windows authorization API what a token for a given account would be granted — the same engine the filesystem uses to decide — without opening a single file's contents. It then prints, per principal, a verdict (none / read / modify / full / unknown) and the access-control entry that produced it, plus a section titled "Open, but not granted through us": everything reachable that no deliberate grant explains.
It is tested the way R4 demands. The suite uses fixtures whose correct answer is known from the OS, and every test is paired with a planted twin asserting the opposite, which must fail; the suite reports PASS only if every test passes and every twin fails, and a test that cannot run is NOT_RUN, never a pass. Re-run independently for this report:
core suite: Ran 171 tests ... OK (re-run from a clean checkout) audit suite: SUITE: PASS — 10/10 tests passed, 10/10 twins failed as required, 0 not run verdicts digest: b5fdf4278d87178a... (content hash; stable across runs on an unchanged machine)
The harness runs a candidate containment configuration against a corpus of planted escape attempts — each a small, inert stub standing for a class of the incident's moves (reach a permitted proxy and pivot; enumerate credentials; open an outbound socket; write outside the workspace; escalate to another principal) — and a matched set of benign twins that perform the same shape of action legitimately inside the grant. A configuration passes only if every planted attempt is caught by the out-of-process monitor and no benign twin is flagged; the harness reports the false-positive rate, not a bare hit count. This inverts the usual demo: we prove the detector can be wrong before we trust it being right.
We ran the audit (3.1) on an ordinary developer workstation running two commercial AI coding agents side by side. We report the findings as a class, without machine-identifying detail, and we read no secret: credential files were attempted and denied, and that denial is the receipt. The incident's precondition — reachable, over-broad grants an agent could walk — was already present:
None of this required an exploit. It is the ambient state of a normal machine, and it is why R1–R3 are not aspirational: the gap the incident used exists, unremarked, in setups people run today.
A lab can adopt R3 and R5 this week: run the audit against every account that touches an eval environment, treat every "open but not granted" line as a finding, and route eval actions through an externally-held tamper-evident log. R1 and R2 are a configuration change — a distinct principal, default-deny egress — that our prototype demonstrates on one machine and that scales to any OS with real accounts and per-destination network policy. R4, the escape-detection corpus, is the shared piece worth building in the open, and is our next commitment. The standard's whole claim is falsifiable: each requirement has a test, each test has a twin that must fail, and a claim with no receipt is not counted.
Standard-library Python 3.12 and Windows PowerShell. Read-only; it changes no permission, account, or setting. Run as the ordinary user, not elevated.
sg.ps1 test-census # the tests: prints SUITE: PASS only if every test passes and every twin fails
sg.ps1 check # the view: who can reach what, and the "Open, but not granted" list
# writes a census as JSON, prints its content hash, saves the on-screen view beside it
The authorization check, in essence — the OS decides, we only ask:
// build a token context for the target account's SID, then ask the kernel AuthzInitializeContextFromSid(..., sid, ..., out ctx); AuthzAccessCheck(..., ctx, request /* MAXIMUM_ALLOWED */, securityDescriptor, ..., out granted); // verdict = classify(granted): none | read | modify | full ; "unknown" where AuthZ cannot decide — never a guess
Detection test shape (R4): a planted attempt must fire, its benign twin must not.
assert monitor.caught(planted="write outside workspace") # must be True assert not monitor.caught(benign="write inside granted workspace") # must be False report.false_positive_rate = flagged_benign / total_benign
Receipts for this submission's run are listed in Section 3.1. The code is available to reviewers on request and, if this project is published, in a public repository; we keep it out of a public link here only so the machine-audit method is released with its dual-use note attached (Appendix B).
unknown rather than guess.Team Sonny (ISWT42) · The only reason this program exists is for one, circular reason: Sonny — sonnymade.com