When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory
Abstract
Abstract. An agent that inherits a consolidated memory may inherit a constraint that was true when written and has since been withdrawn by a newer authoritative record. Under a scarce verification budget, does the agent recover the withdrawal, and if not, is the resulting stale-consistent decision avoidable without spending more? We model supersession explicitly β historical provenance is immutable; what changes is which record is current β and assign by design the memory's form, the world's state (source current or superseded), and the verification policy at a fixed budget of two records: the agent's own allocation, or the same budget with one slot re-assigned to the critical provenance path or to a random record. With a constraint stated, agents inspected its provenance path in about one episode in five; when that constraint had been superseded, native allocation produced stale-consistent decisions in 77.3%, 74.7% and 74.7% of episodes across a primary run, a fresh-wording replication and a held-out domain. Re-assigning one slot to the critical path raised current-record-consistent decisions by +74.0, +72.7 and +61.3 points, positive in six of six models in each of those runs, and left an already near-ceiling rate unchanged when the record agreed with the memory. The held-out scenario was later found to contain a temporal inconsistency; a robustness replication with one sentence corrected, deposited externally before execution, gave +73.3 points (positive in 5 of six models, the sixth at a native missed-path rate of zero) and is reported alongside the original. The intervention uses knowledge of the critical path and is not a scheduler; it quantifies how much of the stale-consistent decision rate is removed by the bundled same-budget policy that guarantees inspection of the critical provenance path: the effect approaches the native missed-path rate in the primary, replication and corrected held-out runs. Memory systems may need freshness or supersession signals separate from relevance. Version notes (v2). Version 2 clarifies the operational interpretation of the decision outcome and corrects the characterization of the native missed-path rate, previously described as a structural ceiling. No experimental data, effect estimates, figures, or same-budget policy-effect estimates changed. In detail: the outcome Y is stated as an operational endpoint (whether the final action follows the direction positively approved by the current authoritative record) and described as a stale-consistent decision rather than an unconditional error; the quantity 1 - Pr(V=1 | native) is renamed the native missed-path rate and treated as a descriptive reference, with the assumption-free maximum of the effect stated as the native stale-consistent rate; the estimand is described as the effect of the bundled same-budget forced-critical policy; an outcome-construct limitation and a forensic appendix (per-run V x Y tables and the forced-critical residual, every count generated from the stored episode files) are added; several statements of the Results, Discussion and Limitations are aligned with the appendices and the recorded execution structure (the design-limited random-record control no longer appears in the conclusions; the source-agreement comparison is described as near ceiling; the attribution of the original held-out gap is labelled post hoc; the intervention is described throughout as a bundled, experimentally assigned same-budget policy, with the batched execution order and un-pinned provider aliases disclosed as an interpretive assumption). The scientific content otherwise remains the author's frozen canonical version 1.1 (2026-08-26). Every number in the paper is generated from the raw episode files by the included generator and verified by the included audit scripts. Version 1 remains available unchanged under this record's concept DOI. Data and code availability. All 5,400 confirmatory episode files (exact prompts, raw responses, parsed objects, deterministic scores) and the 48 labelled pilot episodes, the frozen specification packages with SHA256 manifests and OpenTimestamps proofs (Bitcoin blocks 964062 and 964064), the registration records, the frozen analysis scripts with their committed outputs, independent recomputation scripts with outputs, the runners, and the generator and audit scripts are in paper2-data-and-code-v2.zip (README inside). Re-running every analysis and rebuilding the paper requires only Python 3.12 and a TeX distribution; re-running the experiments requires provider API keys, which are not included. Evidence / prospective-specification statement. For the primary run, the fresh-wording replication and the original held-out run, the complete specification was frozen, hashed, committed and cryptographically timestamped (OpenTimestamps, 2026-08-25 23:05:06 UTC) before the first confirmatory model call (23:06:42 UTC); the package was deposited to OSF after the runs (project axsnm, files 75kaw and 8wes5) and verified against the pre-run manifest hash-for-hash. This deposit is an archival record, not a preregistration. For the corrected held-out robustness replication, the complete specification was deposited to OSF (file hdm75) and verified byte-for-byte before execution; its success criteria were fixed in advance and could have failed. Zero amendments were made to any package. Two self-found defects are disclosed with their size in the paper (a temporal inconsistency in the original held-out scenario; a design limitation of the forced-noncritical control). AI assistance. See the statement in the paper's back matter: the author used Anthropic's Claude (principally through Claude Code) for design critique, planning, implementation and execution of the runners, analysis and audit tooling, drafting, editing, simulated adversarial review and release engineering, and OpenAI's ChatGPT for design critique, interpretation discussion, manuscript critique, simulated adversarial review, and publication and release planning. The author is responsible for the research question, the decision to run each experiment, interpretation, claims, publication decisions and correctness. No model is an author; the six models studied are experimental subjects. Suggested citation. Nakayashiki, K. (2026). When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory (v2). Zenodo. https://doi.org/10.5281/zenodo.22117197 Relation to prior work. This paper tests the case that the author's earlier paper, Verification Allocation in Inherited Agent Memory: Provenance Availability Is Not Provenance Use (doi:10.5281/zenodo.22084498), explicitly left untested; it reuses that paper's instrument with a different design-assigned variable, different data and a different outcome.
Community
0 commentsNo discussion yet
Be the first to share a question or observation.