Synthetic vs. Real Benchmarks for Smart Contract Vulnerability Severity Prediction
Abstract
Context. Machine learning approaches for smart contract vulnerability detection are typically evaluated on synthetic benchmarks of programmatically generated code snippets. Practitioner reports and recent independent evaluations indicate that automated tools continue to miss critical vulnerabilities in production audits, yet the contribution of benchmark selection to this gap remains under-examined.Objectives. This study investigates whether surface-level Solidity features that correlate with vulnerability severity in synthetic benchmarks retain their predictive validity on professionally audited contracts, and proposes a quantitative metric for assessing benchmark suitability for severity prediction research.Methods. Fifteen features were extracted identically from a 10,448-sample synthetic Solidity benchmark and DAppSCAN, a corpus of 1,646 findings from 1,199 audit reports authored by 29 firms. Feature-severity correlations were compared using Fisher r-to-z, Kolmogorov-Smirnov, and Levene tests. Logistic Regression, Random Forest, and Gradient Boosted classifiers were trained in three conditions: in-distribution synthetic, in-distribution real, and cross-distribution.Results. Mean absolute correlation was 0.228 on synthetic data versus 0.057 on real data, a fourfold gap (all p
Community
0 commentsNo discussion yet
Be the first to share a question or observation.