Machine learning reduces audit detection risk in 3.3 million public sector general ledger transactions
Abstract
Abstract Machine learning models for audit anomaly detection are commonly evaluated using proprietary or synthetic datasets, with limited validation against official audit outcomes. This study proposes a four-layer ML framework designed to reduce audit detection risk and evaluates its performance on 3,329,189 general ledger transactions from three consecutive fiscal years (FY2023 to FY2025) of a Mongolian public sector energy utility. The dataset comprises 9,909 account rows and a cumulative debit flow of MNT 16.34 trillion. The proposed framework integrates unsupervised ensemble labeling through Isolation Forest, Z-score analysis, and debit-credit ratio screening, followed by supervised classification with Random Forest, Gradient Boosting, and Decision Tree models. An explainable AI layer maps SHAP feature attributions to specific ISA requirements. Against a simulated 20% MUS baseline, Random Forest achieves F1 = 0.966, AUC = 0.999, and Detection Risk = 2.01%, compared to MUS Detection Risk of 38.05% to 52.7%. McNemar’s test confirms statistically significant superiority (χ² = 1,666.63, p
Community
0 commentsNo discussion yet
Be the first to share a question or observation.