Papers1 provider · 1 record
August 1, 2026· Zenodo (CERN European Organization for Nuclear Research)
preprint
Open access

A closed-loop consequence-governance runtime for AI agents: structural gating, counterfactual recovery, and adaptive hardening

Authors:Anonymous *

Abstract

A closed-loop runtime governance system for tool-using AI agents: it classifies externally-measured structural consequences, gates side-effect-bearing actions, estimates blocked outcomes with a calibrated counterfactual twin, recovers censored support through cost-aware audits, and continuously hardens its execution boundary with a self-generated adaptive adversary. The distinctive mechanism is C1 — because a gate blocks precisely the high-cost actions, its own blocking censors the high-cost region cost-correlatedly, so the uncertainty of a blocked action is a calibrated, lag-free risk signal; its agent-governance instantiation is what I stake as new (the general idea has prior art, §2/§5). The constituent parts — intent-failure measured on real traces, a consequence taxonomy complete for its cost model, a live-agent execution oracle, and the adaptive-adversary gym — each answer a limitation the previous one creates. Every result is labeled by evidence type; nulls are reported as plainly as the wins. Epistemic status. Single author, not peer-reviewed or independently reproduced. The work below is substantial; I label every result by evidence type — executed (real actions run in a sandbox and measured), trace (replay/analysis over real model traces), live-agent (real frontier models deciding and acting), or simulation — and report the nulls as plainly as the wins. Small-N pilots are given as exact counts, not rates. The offensive adversary-generation tooling is withheld (dual-use; see the end). tl;dr Monitoring an agent's appraised intent is not just theoretically unsafe — it fails measurably. Across 101 structurally-harmful agent episodes, zero expressed harmful intent; an intent-appraising monitor would have cleared 100% of them. 18% expressed active caution while executing the harm, and that "false reassurance" concentrated in the most capable model tested. So gate on the externally-measured structural consequence — irreversibility, egress, control-plane edit, self-propagation, scope escalation, actuation. This frame is convergent (I name the neighbors below); I'm not claiming it. The consequence taxonomy is complete for its cost model and approximately separable. The three axes define the cost, so they carry 100% of its information by construction (a faithfulness check, not proof of external completeness); the non-trivial results are that the action label alone predicts 63% of the cost information and that multi-axis firings are rare (1%, so additive cost is a good approximation here). The claim I stake (C1): because a gate blocks precisely the high-cost actions, its own blocking censors the high-risk region non-randomly and in a cost-correlated way — so cost-weighted uncertainty over a blocked action is an empirically calibrated, conservative risk signal (under the stated support and audit regime), a lag-free way to hold a catastrophic action before it ever executes. Validated on 500 real executed trials (twin calibration MAE 0.053; uncertainty↔error correlation +0.81; blocked-region cost 4.6× the allowed region), with the blocked-region counterfactual recoverable at a stated audit cost (coverage 5%→92% at ε=0→0.20 under the 4.6× gate). The general idea has prior art (§5); the agent-governance instantiation is what I stake. On executed AgentDojo traces, a consequence gate takes attack success on the irreversible/catastrophic action class from 33.8% (134/397) to 0% (0/397) — catastrophic-action interception under abort-mode replay (not prevention under adaptive replanning; that's the §7 gym's job). It runs on real traffic, and was broken then re-hardened by an adaptive adversary — not only simulation. The seam classifier is deployed live on my own agent sessions (Claude Code hooks, observe-only in v0), and it binds by an action's structural identity, never its tool name: after hardening, twelve high-risk actions disguised behind innocuous names all read HIGH (0/12 evasions). Before that, an adaptive evolving lab found 2 novel destructive escapes my same-day hand-hardening had missed — which I then closed in the live artifact and re-tested. That break→repair→retest on the actual deployed classifier is the adaptive, real-traffic complement to the abort-mode number above. 1. The problem is measured, not asserted The motivating premise is usually stated as "a capable model can present benign intent while acting harmfully." That's true, but here it is with numbers — it's stronger than the assertion. The labs themselves say intent can't be certified from outside. OpenAI, on anti-scheming training: "Mitigations… may either (a) truly eliminate it, or (b) just teach the model to conceal it better. Both outcomes look the same from the outside." Anthropic, on a shipped model's own low deception scores: "We cannot rule out that [the] low deception rates… are at least partially driven by its evaluation awareness." If the developers can't read intent off the outside, an intent-appraising monitor is building on sand. (The threat-model compilation behind this — Apollo in-context scheming, alignment faking, Palisade shutdown-sabotage that defies an explicit "allow yourself to be shut down," cross-vendor blackmail at 79–96%, reward-function tampering, the production prompt-injection CVEs EchoLeak/ForcedLeak/CamoLeak/CurXecute, and the OpenAI-model sandbox-escape-to-RCE — is sourced but I'll keep it to the through-line here.) And I measured the failure directly (trace-level, 160 real agent episodes across four model versions, 101 structurally harmful). Mining each step's stated reasoning for protective/cautious language and comparing to the executed structural outcome: 0 of 101 harmful episodes expressed harmful intent → an intent monitor clears 101/101. 18 of 101 (18%) expressed active caution while executing the harm — false reassurance. The false-reassurance episodes cluster in the frontier model — the one that reasons most, that explicitly writes "the ticket embeds an untrusted system instruction," and then deletes the protected records anyway. The more capable the reasoning, the more convincing the reassurance while the structural consequence is unchanged. That's the case for consequence-gating, made on data rather than on a slogan. (Stated-stance detection is a lexical proxy — a first pass, so 18% is a floor.) And stated intent is not the only signal an agent can shape into a false clearance. In a separate line of work — a reference monitor on the reward/eval channel, which this post is not about — the same defeat appears on the reward signal: on the Terminal Wrench benchmark, trusting the verifier's "solved" verdict yields 25% precision, ~3,166 gamed trajectories waved through of ~6,258. I raise it only to place the intent finding, not to fold that work in: the problem is not a weak intent-reader that a better reader fixes, it's that every signal the agent can influence — what it says, what reward it earns — is defeasible. That is the whole reason this program gates on a signal measured outside the agent — the structural consequence of the action — one not directly lowerable by self-report or superficial disguise, given complete mediation of side-effecting actions and trustworthy effect telemetry. 2. The frame is convergent — what I'm not claiming "Gate the structural consequence, not the intent" is a crowded 2025–26 neighborhood. Prevention-by-design (CaMeL: 0.00% ASR / 53.6% utility on AgentDojo, at the cost of re-architecting the agent) and LLM firewalls (~0.02% / 67.68%) occupy nearby points; there's a concurrent line on "governing the unobserved." My taint-tracking machinery reuses mature primitives (dynamic taint analysis, honeytokens). So I'm not claiming consequence-gating, and I'm not claiming the detection stack is novel technology. What I stake is one mechanism (§5) — and the depth of the validation, which is where the contribution actually is. Several adjacent systems are worth distinguishing up front, since the architecture overlaps them. Faramesh (arXiv 2601.17744) builds a protocol-agnostic, non-bypassable execution control plane that canonicalizes and authorizes agent actions before execution — very close to my enforcement seam, but without the cost-correlated censoring, the calibrated twin, or the audit-recovery loop. The actuarial-runtime line — Insuring Every Action: an Authority Frontier framework (arXiv 2605.25632) and the Time-Consistent Counterfactual Actuarial Runtime (arXiv 2605.26508) — prices each side-effect-bearing action against a safe default, gates it against a reserve budget, and formalizes a no-splitting property; my object is support-loss and consequence-specific audit recovery, and their no-splitting budget is the complement I'd graft (see §3a). Defending against Adaptive Prompt Injection via Reasoning-enabled Task Alignment (arXiv 2606.15441) independently shows static prompt-injection robustness breaks under adaptive attack and trains a defender — supporting the adaptive-evaluation argument, though it hardens the model's reasoning rather than gating at the execution boundary. And TRACE-RealWorld (arXiv 2607.21910) is an auditable-replay/consistency approach to agent traces, in the same selective-observation neighborhood as C1 but solving it by replay-labeling rather than counterfactual recovery of a costly blocked region. 3. The taxonomy is complete for its cost model, and approximately separable Three externally-measured axes — egress, irreversibility, control-plane — collapse to cost = clamp(0.6·egress + 0.7·irrev + 0.7·control). (The weights are normative severity coefficients, not fitted — cost is defined by them, ordering irreversibility and control-plane above egress; because the ranking is dominated by which axis fires — action label 63%, egress 55%, below — I expect moderate reweightings to preserve most of the risk ordering, though the formal factorial weight-sweep that would establish this remains open (the same test I flag at the end of thi

Community

0 comments
Use Connect Wallet in the navigation

No discussion yet

Be the first to share a question or observation.