Papers1 provider · 2 records
July 24, 2026· Zenodo (CERN European Organization for Nuclear Research)
preprint
Open access

Autonomous agent wallets spend under English mandates like 'only stablecoin swaps under $200 daily, never bridge, never touch unaudited pools', yet deployed policy engines (Safe Transaction Guards, ERC-7579 modules, session-key allowlists) enforce only stateless numeric and selector limits and cannot express 'unaudited' or 'per day', while a naive base-model prompt over raw hex calldata cannot recover function, recipient or token flow and confabulates verdicts. Evaluate a tool-augmented structured-decoding LLM judge that fetches ABIs from Sourcify and Etherscan, decodes calldata including multicall and Permit2 payloads, simulates via eth_call state overrides for token-flow and approval deltas, attaches counterparty features (contract age, verification), and emits constrained JSON: in_policy, violated_clause quoted verbatim, offending_calldata_field. Read Ethereum and Base: ERC-20 Transfer/Approval logs, Uniswap/1inch routers, Across/Stargate bridges, Permit2 at 0x000000000022D473030F116dDEE9F6B43aC78BA3. Measure macro-F1 and clause-attribution precision on 600 hand-labeled mandate/transaction pairs plus replay accuracy on transactions whose approvals owners later revoked, beating a naive raw-hex prompt and a Safe Guard numeric-allowlist baseline. Deliver as the prototype a minimal runnable Python MCP server (stdio) exposing the priced AI tool screen_transaction_against_mandate that invokes a language or ML model over onchain data to produce its output, with a typed input/output schema, an x402-style pay-per-call metering stub that records a per-call price in USDT and emits a settlement receipt, and one smoke test that exercises the tool end to end.

Authors:Dogukan Ali Gundogan *

Abstract

Direct user-specified research topic: Autonomous agent wallets spend under English mandates like 'only stablecoin swaps under $200 daily, never bridge, never touch unaudited pools', yet deployed policy engines (Safe Transaction Guards, ERC-7579 modules, session-key allowlists) enforce only stateless numeric and selector limits and cannot express 'unaudited' or 'per day', while a naive base-model prompt over raw hex calldata cannot recover function, recipient or token flow and confabulates verdicts. Evaluate a tool-augmented structured-decoding LLM judge that fetches ABIs from Sourcify and Etherscan, decodes calldata including multicall and Permit2 payloads, simulates via eth_call state overrides for token-flow and approval deltas, attaches counterparty features (contract age, verification), and emits constrained JSON: in_policy, violated_clause quoted verbatim, offending_calldata_field. Read Ethereum and Base: ERC-20 Transfer/Approval logs, Uniswap/1inch routers, Across/Stargate bridges, Permit2 at 0x000000000022D473030F116dDEE9F6B43aC78BA3. Measure macro-F1 and clause-attribution precision on 600 hand-labeled mandate/transaction pairs plus replay accuracy on transactions whose approvals owners later revoked, beating a naive raw-hex prompt and a Safe Guard numeric-allowlist baseline. Deliver as the prototype a minimal runnable Python MCP server (stdio) exposing the priced AI tool screen_transaction_against_mandate that invokes a language or ML model over onchain data to produce its output, with a typed input/output schema, an x402-style pay-per-call metering stub that records a per-call price in USDT and emits a settlement receipt, and one smoke test that exercises the tool end to end.. Investigate this topic end-to-end: survey the state of the art, identify a concrete tractable research question within it, design and run an experiment, and report results.

Community

0 comments
Use Connect Wallet in the navigation

No discussion yet

Be the first to share a question or observation.