Temporal Dynamics of Memory Poisoning in Web3-Style LLM Agents
Abstract
Memory-enabled large language model (LLM) agents, particularly those deployed in long-horizon, tool-using settings such as Web3-style autonomous workflows, introduce security risks that extend beyond single-prompt injection. By persisting and reusing information across interaction steps and sessions, these agents enable memory poisoning attacks in which adversarial inputs modify persistent agent state and influence future decisions after benign intermediate interactions. Recent work on context manipulation and “fake memories” demonstrates that adversarial content can be injected into an agent’s prompt-visible inputs or persistent memory; however, existing evaluations largely analyze such attacks at isolated interaction steps or static context snapshots, obscuring their temporal dynamics. In this paper, we present the first large-scale, trajectory-level measurement framework for analyzing temporal memory poisoning in memory-enabled LLM agents. We construct a schema-constrained dataset of 2,614 multi-step attack trajectories spanning four attack families,chain poisoning, policy rewriting, backdoor triggering, andslow drift, executed over shared persistent memory. We define temporal risk metrics over multi-step interaction trajectories that capture delayed activation, non-monotonic escalation, and the earliest point at which attacks become distinguishable from benign behavior. Our empirical results show that a substantial fraction of attacks remain indistinguishable from benign behavior until late-stage activation, despite exhibiting low or medium risk at all earlier steps. Slow-drift and backdoor-trigger attacks, in particular, systematically evade step-local evaluation until terminal interactions, while chain poisoning and policy rewriting exhibit non-monotonic risk trajectories. These findings demonstrate that memory poisoning risk is inherently temporal and cannot be reliably assessed using prompt-level or step-isolated evaluation, motivating trajectory-aware benchmarks for agent security.
Community
0 commentsNo discussion yet
Be the first to share a question or observation.