Decentralized Finance (DeFi), a financial ecosystem without centralized controlling organization, has introduced a new paradigm for lending and borrowing. However, its capital efficiency remains constrained by the inability to effectively assess the risk associated with each user/wallet. This paper introduces the 'On-Chain Credit Risk Score (OCCR Score) in DeFi', a probabilistic measure designed to quantify the credit risk associated with a wallet. By analyzing historical real-time on-chain activity as well as predictive scenarios, the OCCR Score may enable DeFi lending protocols to dynamically adjust Loan-to-Value (LTV) ratios and Liquidation Thresholds (LT) based on the risk profile of a wallet. Unlike existing wallet risk scoring models, which rely on heuristic-based evaluations, the OCCR Score offers a more objective and probabilistic approach, aligning closer to traditional credit risk assessment methodologies. This framework can further enhance DeFi's capital efficiency by incentivizing responsible borrowing behavior and optimizing risk-adjusted returns for lenders.
Gibran Gómez, Kevin van Liebergen, Davide Sanvito, Giuseppe Siracusano · 6 authors
Cryptocurrency abuse reporting services are a valuable data source about abusive blockchain addresses, prevalent types of cryptocurrency abuse, and their financial impact on victims. However, they may suffer data pollution due to their crowd-sourced nature. This work analyzes the extent and impact of data pollution in cryptocurrency abuse reporting services and proposes a novel LLM-based defense to address the pollution. We collect 289K abuse reports submitted over 6 years to two popular services and use them to answer three research questions. RQ1 analyzes the extent and impact of pollution. We show that spam reports will eventually flood unchecked abuse reporting services, with BitcoinAbuse receiving 75% of spam before stopping operations. We build a public dataset of 19,443 abuse reports labeled with 19 popular abuse types and use it to reveal the inaccuracy of user-reported abuse types. We identified 91 (0.1%) benign addresses reported, responsible for 60% of all the received funds. RQ2 examines whether we can automate identifying valid reports and their classification into abuse types. We propose an unsupervised LLM-based classifier that achieves an F1 score of 0.95 when classifying reports, an F1 of 0.89 when classifying out-of-distribution data, and an F1 of 0.99 when identifying spam reports. Our unsupervised LLM-based classifier clearly outperforms two baselines: a supervised classifier and a naive usage of the LLM. Finally, RQ3 demonstrates the usefulness of our LLM-based classifier for quantifying the financial impact of different cryptocurrency abuse types. We show that victim-reported losses heavily underestimate cybercriminal revenue by estimating a 29 times higher revenue from deposit transactions. We identified that investment scams have the highest financial impact and that extortions have lower conversion rates but compensate for them with massive email campaigns.
Abstract This study provides a comprehensive review of machine learning (ML) applications in the fields of business and finance. First, it introduces the most commonly used ML techniques and explores their diverse applications in marketing, stock analysis, demand forecasting, and energy marketing. In particular, this review critically analyzes over 100 articles and reveals a strong inclination toward deep learning techniques, such as deep neural, convolutional neural, and recurrent neural networks, which have garnered immense popularity in financial contexts owing to their remarkable performance. This review shows that ML techniques, particularly deep learning, demonstrate substantial potential for enhancing business decision-making processes and achieving more accurate and efficient predictions of financial outcomes. In particular, ML techniques exhibit promising research prospects in cryptocurrencies, financial crime detection, and marketing, underscoring the extensive opportunities in these areas. However, some limitations regarding ML applications in the business and finance domains remain, including issues related to linguistic information processes, interpretability, data quality, generalization, and the oversights related to social networks and causal relationships. Thus, addressing these challenges is a promising avenue for future research.
Related party transactions (RPTs) can serve as channels for the spread of credit risk events among blockchain firms. However, current credit risk-assessment models typically only consider a firm’s individual characteristics, overlooking the impact of related parties in the blockchain. We suggest incorporating RPT network analysis to improve credit risk evaluation. Our approach begins by representing an RPT network using a weighted adjacency matrix. We then apply DANE, a deep network embedding algorithm, to generate condensed vector representations of the firms within the network. These representations are subsequently used as inputs for credit risk-evaluation models to predict the default distance. Following this, we employ SHAP (Shapley Additive Explanations) to analyze how the network information contributes to the prediction. Lastly, this study demonstrates the enhancing effect of using DANE-based integrated features in credit risk assessment.
As the non-fungible token (NFT) market flourishes, price prediction emerges as a pivotal direction for investors gaining valuable insight to maximize returns. However, existing works suffer from a lack of practical definitions and standardized evaluations, limiting their practical application. Moreover, the influence of users' multi-behaviour transactions that are publicly accessible on NFT price is still not explored and exhibits challenges. In this paper, we address these gaps by presenting a practical and hierarchical problem definition. This approach unifies both collection-level and token-level task and evaluation methods, which cater to varied practical requirements of investors. To further understand the impact of user behaviours on the variation of NFT price, we propose a general wallet profiling framework and develop a COmmunity enhanced Multi-bEhavior Transaction graph model, named COMET. COMET profiles wallets with a comprehensive view and considers the impact of diverse relations and interactions within the NFT ecosystem on NFT price variations, thereby improving prediction performance. Extensive experiments conducted in our deployed system demonstrate the superiority of COMET, underscoring its potential in the insight toolkit for NFT investors.
The development of digital technologies that increase the security of payments and improve settle-ments is one of the main tasks facing regulators in the financial sector. The main reason for the in-creased attention to distributed ledger technology in the financial sector is the expectation that it will eliminate a number of problems and limitations inherent in the currently used methods of storing, ac-counting and transmitting financial information. The authors consider the advantages of the distribut-ed ledger technology in the implementation of financial transactions
The accurate imputation of missing values in time series data is paramount for maintaining the integrity and reliability of analyses and predictions. This article investigates the effica-cy of various missing values imputation methods, encom-passing well-known machine learning and statistical tech-niques. Moreover, for a better understanding, they imple-mented two financial data time series: S&P 500 and Bitcoin markets spanning from 2016 to 2023 on a daily frequency. Initially utilizing complete datasets, controlled missingness was introduced by randomly removing 45 data points. Then, these methods applied multiple imputation strategies for estimating and substituting these missing values. Experi-mental evaluation yielded insightful findings regarding the performance of the different methods. The examined ma-chine learning methods, including k-Nearest Neighbors (k-NN), Random Forest, Deep Learning, and Decision Trees, consistently outperformed their statistical counterparts, such as Mean Imputation, Regression Imputation, Hot-Deck Im-putation, and Expectation-Maximization Imputation. Nota-bly, Random Forest emerged as the most effective method, showcasing superior performance in terms of accuracy and robustness. Conversely, the Mean Imputation method exhibited com-paratively inferior outcomes, suggesting its limited suitabil-ity for financial time series data. This research contributes to the ongoing discourse on data integrity within finance ana-lytics and serves as a comprehensive guide for practitioners seeking optimal missing values imputation methods. The empirical evidence provided herein advances the under-standing of imputation techniques' relative performance and their application in financial data, facilitating enhanced de-cision-making processes and yielding more reliable predic-tions.
Salman Bahoo, Marco Cucculelli, Xhoana Goga, Jasmine Mondolo
Abstract Over the past two decades, artificial intelligence (AI) has experienced rapid development and is being used in a wide range of sectors and activities, including finance. In the meantime, a growing and heterogeneous strand of literature has explored the use of AI in finance. The aim of this study is to provide a comprehensive overview of the existing research on this topic and to identify which research directions need further investigation. Accordingly, using the tools of bibliometric analysis and content analysis, we examined a large number of articles published between 1992 and March 2021. We find that the literature on this topic has expanded considerably since the beginning of the XXI century, covering a variety of countries and different AI applications in finance, amongst which Predictive/forecasting systems, Classification/detection/early warning systems and Big data Analytics/Data mining /Text mining stand out. Furthermore, we show that the selected articles fall into ten main research streams, in which AI is applied to the stock market, trading models, volatility forecasting, portfolio management, performance, risk and default evaluation, cryptocurrencies, derivatives, credit risk in banks, investor sentiment analysis and foreign exchange management, respectively. Future research should seek to address the partially unanswered research questions and improve our understanding of the impact of recent disruptive technological developments on finance.
Marjan Alirezaie, William Hoffman, Paria Zabihi, Hossein Rahnama · 5 authors
The complexities arising from disparate data sources, conflicting contracts, residency requirements, and the demand for multiple AI models in trade finance supply chains have hindered small and medium-sized enterprises (SMEs) with limited resources from harnessing the benefits of artificial intelligence (AI) capabilities, which could otherwise enhance their business efficiency and predictability. This paper introduces a decentralized AI orchestration framework that prioritizes transparency and explainability, offering valuable insights to funders, such as banks, and aiding them in overcoming the challenges associated with assessing SMEs’ financial credibility. By utilizing an orchestration technique involving symbolic reasoners, language models, and data-driven predictive tools, the framework empowers funders to make more informed decisions regarding cash flow prediction, finance rate optimization, and ecosystem risk assessment, ultimately facilitating improved access to pre-shipment trade finance for SMEs and enhancing overall supply chain operations.
Emerging blockchain payment gateways have facilitated worldwide financial systems with unprecedented efficiency, transparency, and decentralization. Yet, increasingly, such platforms become susceptible to complex financial risks such as fraud at various scales, double-spending, Sybil attacks, and illegal access. The rule-based approaches that were traditionally implemented are no longer adequate to keep up with the evolving threat landscape of decentralized finance (DeFi). This, therefore, serves to strengthen the stance for considering ML models for at least real-time transaction analysis and fraud detection. Even with many models offering good prediction capabilities, the lack of transparency raises serious concerns about issues of interpretability and compliance—especially in environments that are financially regulated. The paper thus delves into the incorporation of explainable machine learning (XML) techniques in blockchain payment risk assessment frameworks. Using model-agnostic tools such as SHAP (SHapley Additive Explanations) and LIME (Local Interpretable Model-Agnostic Explanations) and combining them with very high-end models such as XGBoost and LightGBM, we create interpretable frameworks that enable stakeholders to understand, trust, and verify the risk classifications issued. Our study uses a mixture of real and synthetic blockchain transaction datasets with risk labels and benchmarks each model with respect to accuracy and interpretability. Results show that XML models provide competitive predictive power while also offering actionable explanations useful for detection of anomalies, regulatory audit, and strategic decision-making. We believe that explainable ML is not just achievable but also an absolute prerequisite for sustainable and compliant risk management in blockchain financial infrastructures
This study introduces an interpretable imbalanced data classification method for detecting cryptocurrency transaction fraud. We address data imbalance using SMOTE oversampling and data augmentation through contrastive learning. Next, we introduce a Transformer-based deep learning model that learns sample relevance. The model undergoes pre-training with a contrastive loss and fine-tuning through Bayesian optimization to effectively extract high-dimensional, higher-order, and fraud-related features. We employ a SHAP-based interpreter along with attention scores to elucidate the role of various transaction features in fraud detection. Comparative results demonstrate the model's remarkable recall performance in identifying cryptocurrency transaction fraud. Furthermore, it achieves an excellent F1 value, striking a balance between accuracy and recall. This research not only enriches financial fraud detection but also enhances cryptocurrency transaction security, promotes market development, and contributes to economic stability and social security.
C. Vinoth Kumar, Poongundran Selvaprabhu, Nivetha Baska, Vivek Menon U · 7 authors
The Know Your Customer (KYC) process is a fundamental prerequisite for any financial institution’s compliance with the regulatory framework. Blockchain technology has emerged as a revolutionary solution to enhance the effectiveness of the KYC procedure. It ensures that the KYC process is transparent, secure, and immutable, thereby offering a robust solution to combat fraudulent activities. The potential of blockchain technology in revolutionizing the KYC process has been acknowledged globally. Blockchain technology provides a decentralized platform for storing customer data, enabling financial institutions to access the information seamlessly. Using ethereum blockchain technology in KYC procedures can enhance the efficiency of financial institutions, significantly reducing the time and cost associated with the process. This work aims to provide a viable and sustainable solution to the challenges that banks experience in implementing KYC procedures and onboarding new customers. The proposed solution involves the central bank maintaining a comprehensive register of all registered banks while closely monitoring their adherence to the existing regulations governing KYC and customer acquisition.
Collateral is an item of value serving as security for the repayment of a loan. In blockchain-based loans, cryptocurrencies serve as the collateral. The high volatility of cryptocurrencies implies a serious barrier of entry with a common practice that collateral values equal multiple times the value of the loan. As assets serving as collateral are locked, this requirement prevents many candidates from obtaining loans. In this paper, we aim to make loans more accessible by offering loans with lower collateral, while keeping the risk for lenders bound. We use a credit score based on data recovered from the blockchain to predict how likely someone is to repay a loan. Our protocol does not risk the initial amount granted by liquidity providers, but only risks part of the interest yield gained by the protocol in the past.
Olawole Akomolafe, Babajide Oluwaseun Olaogun, Michael Olumuyiwa Adesuyi, Victor Ukara Ndukwe · 5 authors
Effective liquidity management is critical for the reliability and efficiency of international remittance and cross-border payment systems. Delays, settlement failures, and currency conversion inefficiencies can significantly impact SMEs, corporates, and individual remitters, leading to operational disruptions, increased costs, and reduced financial inclusion. This study proposes a Predictive AI Model for Remittance Liquidity Optimization, designed to forecast liquidity requirements in real time, optimize fund allocation, and enhance the overall performance of international payment networks. The model integrates multi-source data, including historical transaction volumes, foreign exchange (FX) rates, settlement schedules, and network congestion metrics, to generate predictive insights and automated liquidity management recommendations. The conceptual framework of the model incorporates advanced machine learning and time-series forecasting techniques, combined with an optimization engine that dynamically allocates available funds to minimize delays, reduce transaction costs, and manage FX risks. Real-time anomaly detection mechanisms identify potential liquidity shortfalls, network congestion, or settlement failures, triggering alerts and corrective actions. The model also includes integration layers with banking platforms, fintech providers, and remittance networks, enabling seamless execution of liquidity redistribution and settlement optimization. Predictive outputs are visualized through interactive dashboards, supporting operators in decision-making and ensuring transparency in fund flows. By leveraging AI-driven forecasting and optimization, the model reduces settlement delays, improves FX efficiency, and enhances operational reliability across multi-currency, multi-jurisdictional payment corridors. Its applications extend to SMEs, corporate treasuries, and high-volume remittance corridors, promoting financial inclusion and operational continuity. Future extensions include adaptive learning algorithms for self-optimizing liquidity strategies, integration with distributed ledger technologies for real-time settlements, and expansion to multi-party global supply chains. Ultimately, this predictive AI model provides a scalable, intelligent solution for enhancing liquidity management in international payment systems, fostering greater efficiency, resilience, and transparency in global financial networks.
As one of the most popular blockchain platforms supporting smart contracts, Ethereum has caught the interest of both investors and criminals. Differently from traditional financial scenarios, executing Know Your Customer verification on Ethereum is rather difficult due to the pseudonymous nature of the blockchain. Fortunately, as the transaction records stored in the Ethereum blockchain are publicly accessible, we can understand the behavior of accounts or detect illicit activities via transaction mining. Existing risk control techniques have primarily been developed from the perspectives of de-anonymizing address clustering and illicit account classification. However, these techniques cannot be used to ascertain the potential risks for all accounts and are limited by specific heuristic strategies or insufficient label information. These constraints motivate us to seek an effective rating method for quantifying the spread of risk in a transaction network. To the best of our knowledge, we are the first to address the problem of account risk rating on Ethereum by proposing a novel model called RiskProp, which includes a de-anonymous score to measure transaction anonymity and a network propagation mechanism to formulate the relationships between accounts and transactions. We demonstrate the effectiveness of RiskProp in overcoming the limitations of existing models by conducting experiments on real-world datasets from Ethereum. Through case studies on the detected high-risk accounts, we demonstrate that the risk assessment by RiskProp can be used to provide warnings for investors and protect them from possible financial losses, and the superior performance of risk score-based account classification experiments further verifies the effectiveness of our rating method.
Existing works on valuing digital assets on the Internet typically focus on a single asset class. To promote the development of automated valuation techniques, preferably those that are generally applicable to multiple asset classes, we construct DASH, the first Digital Asset Sales History dataset that features multiple digital asset classes spanning from classical to blockchain-based ones. Consisting of 280K transactions of domain names (DASH_DN), email addresses (DASH_EA), and non-fungible token (NFT)-based identifiers (DASH_NFT), such as Ethereum Name Service names, DASH advances the field in several aspects: the subsets DASH_DN, DASH_EA, and DASH_NFT are the largest freely accessible domain name transaction dataset, the only publicly available email address transaction dataset, and the first NFT transaction dataset that focuses on identifiers, respectively. We build strong conventional feature-based models as the baselines for DASH. We next explore deep learning models based on fine-tuning pre-trained language models, which have not yet been explored for digital asset valuation in the previous literature. We find that the vanilla fine-tuned model already performs reasonably well, outperforming all but the best-performing baselines. We further propose improvements to make the model more aware of the time sensitivity of transactions and the popularity of assets. Experimental results show that our improved model consistently outperforms all the other models across all asset classes on DASH.
Fraud prevention in cryptocurrency transactions is paramount because the proportion of fraudulent transactions in this sector is increasing.Artificial intelligence (AI) can increase the ability to identify fraudulent transactions through patterns and irregularities that are hardly noticeable through other approaches.This paper focuses on AI approaches, such as machine learning algorithms, LightGBM, and fraud detection.Some notable works are comparing AI strategies, methods to overcome difficulties when using imbalanced data sets, and the practical utility of these models.
Cryptocurrency networks that provide a new way of securing financial transactions have gained a surge of interest in recent years. However, recent studies reveal that blockchain networks are rampant with frauds and are prone to several privacy and security issues. The public availability of cryptocurrency transaction records provides an unprecedented opportunity for researchers to analyze cryptocurrency transactions. In particular, anomaly detection techniques are promising avenues for fighting illicit activities such as money laundering, terrorist financing, drug trafficking, scams, frauds, and many more.In this dissertation, we address the challenges of detecting anomalous entities on cryptocurrency networks. Firstly, we introduce effective techniques for generating features for network entities that are directly devised from raw data and highlight the utility of those features in detecting illicit accounts on the Ethereum network. Next, we enrich the proposed method by expanding the feature set through the incorporation of graph-based features that embed the relational information of networks. This also enables us to generalize our methods to instances of cryptocurrency networks with different architectural models. Based on the success of our method in anomaly detection in cryptocurrency networks, we further generalize our model to encompass a generic temporal weighted multidigraph and show the state-of-the-art results for anomaly detection in other common domains including rating and social networks. In doing so, we also investigate the challenges of employing node classification techniques for anomaly detection, which is a common practice. Here, we discuss the importance of performance metrics and evaluation settings when interpreting the efficiency of different methods and tasks, which is often overlooked by the community. Finally, we shift our focus to examining the inherent challenges of learning on dynamic networks, which is an important emerging research field with applications in drug discovery, computational finance, social networks, etc. Here, we propose solutions for providing a more robust evaluation setup for dynamic graph learning methods. The key contributions of this dissertation are twofold: First, we describe efficient techniques for detecting anomalies on cryptocurrency networks and generalize them to other real-world complex networks. Second, we focus on the temporal aspect of these networks and investigate how the dynamism of networks affects the downstream tasks and evaluation settings
The increasing digitization of financial services by the late 2010s resulted in the generation of massive volumes of transactional data across payment systems, trading platforms, digital banking applications, and regulatory reporting pipelines. These transaction logs, originally designed for auditing, reconciliation, and failure recovery, gradually emerged as a valuable source of behavioral and operational insight. However, the scale, velocity, and structural heterogeneity of transactional logs posed significant challenges to traditional analytical techniques, which were often optimized for static datasets or narrowly defined reporting use cases. As a result, organizations began exploring systematic approaches to mine patterns from transaction logs in order to better understand system behavior, detect anomalies, and improve decision-making. Pattern mining from transaction logs refers to the process of discovering recurring structures, sequences, correlations, and deviations within recorded transactional events. By September 2019, this practice was informed by a combination of data mining research, distributed systems logging techniques, and operational analytics developed in large-scale production environments. Unlike conventional business intelligence queries, pattern mining emphasizes the identification of latent relationships and temporal structures that are not explicitly encoded in application logic. These patterns may reflect normal operational workflows, emergent system behaviors, or early indicators of faults, fraud, or performance degradation. In financial systems, transaction logs capture more than simple state changes; they encode regulatory-relevant actions such as authorization decisions, settlement progressions, risk evaluations, and ledger mutations. Mining patterns from these logs enables institutions to analyze end-to-end transaction lifecycles, correlate technical events with business outcomes, and identify systemic inefficiencies or vulnerabilities. Importantly, such analysis must operate within strict constraints related to data privacy, auditability, and regulatory compliance, distinguishing transaction log mining in financial domains from analogous practices in less regulated environments. This paper examines pattern mining from transaction logs as understood and applied by September 2019, situating it within the broader evolution of logging, distributed systems observability, and data mining research. It synthesizes academic literature and industry practices to propose a conceptual and architectural framework for extracting meaningful patterns from transactional data at scale. The analysis focuses on methodological considerations, architectural layering, and practical challenges encountered in regulated, high-throughput systems, while avoiding retrospective interpretations based on post-2019 technologies or techniques.
In this paper I discuss how blockchains potentially could affect the way credit risk is modeled, and how the improved trust and timing associated with blockchain-enabled real-time accounting could improve default prediction. To demonstrate the (quite substantial) effect the change would have on well-known credit risk measures, a simple case-study compares Z-scores and Merton distances to default computed using typical accounting data of today to the same risk measures computed under a hypothetical future blockchain regime.