Papers1 provider · 1 record
June 8, 2026· Preprints.org
preprint
Open access

Data Leakage-Free Explainable AI for Decentralized Credit Scoring: A SHAP-Interpretable Approach to Default Prediction

Authors:Sai Srikanth MadugulaPeplluis Esteva De La RosaDaya Shankar

Abstract

The integration of machine learning into decentralized finance (DeFi) credit assessment is frequently undermined by opaque algorithms and severe methodological flaws regarding data leakage. This paper presents a rigorous, fully reproducible framework for explainable artificial intelligence (XAI) in invoice-backed default risk modeling. Utilizing a highly imbalanced dataset of 12,000 corporate loan originations, we engineer an XGBoost ensemble model that achieves an AUC-ROC of 0.89. We systematically eliminate the pervasive data leakage associated with the Synthetic Minority Over-sampling Technique (SMOTE) by implementing a dynamic crossvalidation pipeline, ensuring synthetic data generation is strictly isolated to training folds. To satisfy institutional accounting standards for expected loss (e.g., IFRS 9), we mathematically formulate and validate the Expected Calibration Error (ECE), achieving a highly calibrated probabilistic output of 0.08. Furthermore, we extract local explanations using SHAP (SHapley Additive exPlanations), imposing strict constraints on the background reference dataset to guarantee mathematical additivity and prevent stochastic approximation transitions. Our findings reveal that Days Payment Outstanding (DPO) and invoice age are primary default drivers, while on-chain reputation effectively mitigates perceived risk. Finally, we address critical privacy vulnerabilities, mathematically modeling Membership Inference Attacks (MIAs) on synthetic records. This work establishes a regulatory-compliant, structurally sound ML foundation for permissionless credit provision.

Community

0 comments
Use Connect Wallet in the navigation

No discussion yet

Be the first to share a question or observation.