Allocating Human Judgment in Automated Blockchain Anti-Money Laundering: A Review of Detection, Attribution, Adjudication, and Reporting
Abstract
Automated anti-money laundering (AML) on public blockchains is usually framed as a detection problem. Because the ledger is public and permanent, automated detection is feasible, but that record shows only that value moved, without showing who moved it or why. We review automated blockchain anti-money laundering as a system in which machine models and human analysts share each decision, following it through four stages: detection, attribution, adjudication, and reporting, and asking at each stage what automation does well, what it must leave to human judgment, and what goes wrong when that judgment is misplaced. The evidence shows a consistent asymmetry. Machine learning is strong at pattern-finding over the permanent public record, where models rank suspicion and clustering heuristics scale, but it weakens sharply as the task turns from finding a pattern to assigning meaning, identity, intent, or accountability. Drawing first on emerging AML-specific studies and then, where direct evidence remains insufficient, on human-factors research from adjacent high-stakes domains, we find an uneven evidence base. Automation is best supported for large-scale detection, alert prioritization, and parts of blockchain attribution, while the evidence becomes more limited as decisions require contextual interpretation, evidential judgment, accountability, and reporting. AML-specific studies identify explainability, flexibility, tool integration, supervisory justification, and human review as important operational requirements, but they do not yet establish how frequently analysts over-rely on, reject, or selectively follow automated recommendations. Evidence from aviation, healthcare, and public administration is therefore used to identify plausible failure mechanisms rather than to claim AML-specific effects. The result is an evidence-weighted allocation matrix that records, for each stage, the strength of the case for automation, the function retained by the analyst, the dominant failure mode, and the empirical question that remains unresolved. Support is strongest for detection, direct but context-dependent for attribution, and more provisional for adjudication and reporting.
Community
0 commentsNo discussion yet
Be the first to share a question or observation.