Papers1 provider · 2 records
April 16, 2026· Open MIND
article
Open access

Provable and Practical Prompt Injection Resilience in Autonomous LLM Agents

Authors:Rohith SinghMr. Charan SinghAbdul RashadMd. Abdur RasheedNehith SayiniSomanath Nayak

Abstract

Prompt injection is a foundational security vulnerability in large language models (LLMs) deployed as autonomous agents with tool access and multi-step reasoning capabilities. Existing defenses rely on heuristic filters that fail under obfuscation, indirect injection, and multi-agent propagation. We present a Unified Cryptographic-Control Architecture (UCCA), a principled framework that integrates five complementary guarantees: (1) information-theoretic leakage bounds derived via Fano's inequality, (2) certified robustness via randomized smoothing, (3) token-level rejection via erase-and-check, (4) runtime trajectory enforcement via control barrier functions (CBFs), and (5) verifiable inference via zero-knowledge proofs (ZK-SNARKs). We formally prove that any successful prompt injection attack must simultaneously bypass all five mechanisms, a condition we show has probability at most δ under stated assumptions. We evaluate UCCA on three real LLMs (GPT-4o, Claude 3.5 Sonnet, Mistral-7B) across four established attack benchmarks (INJECAGENT, TensorTrust, PromptBench, HarmBench), achieving attack success rates below 8% while maintaining median latency overhead under 340 ms. Our framework bridges formal security guarantees and deployable system architecture, establishing a foundation for provably secure autonomous AI. • Information-theoretic bounds on system prompt leakage using mutual information and Fano's inequality. • Certified robustness for safety-critical classification through randomized smoothing, where the robustness radius R is determined from output probability gaps. • Token-level rejection guarantees using an erase-and-check procedure capable of detecting adversarial subsets of size ≤ k. • Runtime safety enforcement through control barrier functions (CBFs), ensuring LLM outputs remain within a verified safe set. • Verifiable inference using ZK-SNARKs, allowing cryptographic attestation of model outputs without revealing model weights. • UCCA, a deployable system integrating all five mechanisms, evaluated on real LLMs and standard benchmarks.

Community

0 comments
Use Connect Wallet in the navigation

No discussion yet

Be the first to share a question or observation.