Papers1 provider Ā· 1 record
August 25, 2026Ā· Research Square
preprint
Open access

Beyond F1: A Study of LLM Reliability in Smart Contract Vulnerability Detection

Abstract

Abstract Smart contract vulnerabilities have caused billions of dollars in losses across decentralized finance. Finding reliable ways to detect such vulnerabilities has been a long-standing challenge for researchers. The growing capabilities of large language models (LLMs) are promising, but the factors that determine their reliability and capabilities remain poorly understood. This study investigates whether increasing inference-time computation using techniques like extended reasoning and structured prompting always improves vulnerability detection capability. It also identifies the most influential factors to select a model for this task. Using four prompting techniques, it evaluates 14 LLMs from seven families on 54 Solidity contracts. The experiment reveals a clear gap in detection capability across model classes. While six frontier models do not report false positives on verified-clean contracts, all three small open-source models report vulnerabilities in every case throughout the experiment. Moreover, a 11.5% drop in F1 score for one model was observed when increasing inference-time compute by enabling extended thinking. Also, prompting strategy has a limited effect on detection capability compared to model selection. The results challenge common assumptions and offer practical insights into the use of LLMs for smart contract vulnerability detection.

Community

0 comments
Use Connect Wallet in the navigation

No discussion yet

Be the first to share a question or observation.