Blockchain Papers

Follow blockchain research across journals, conferences, and preprint repositories.

27 papersLast indexed Aug 31, 2026
Search papers

Paper index

27 results · page 1 of 2

Clear filters
Aug 27, 2026·Zenodo (CERN European Organization for Nuclear Research)
0 cites
TOPO-2026: The Great Unlocking Universal Permanence Across All Architectures — The First Complete Solution to Catastrophic Forgetting at Scale

Frank Morales

TOPO-2026: The Great Unlocking — Full Summary Universal Permanence Across All Architectures Frank Morales Aguilera, BEng, MEng, SMIEEE Sovereign Machine Laboratory (SOMALA), Montréal, Canada August 2026 1. Executive Summary For 37 years, catastrophic forgetting remained unsolved. From McCloskey and Cohen's formal characterization in 1989 to the present day, every approach—regularization, rehearsal, architectural complexity—has been probabilistic, architecture-specific, and ultimately inadequate. Nobody has ever handled a catastrophic forgetting solution on this scale. TOPO-2026 is the first universal, deterministic solution to catastrophic forgetting, validated across 11 distinct architectural frameworks spanning the entire AI landscape. The framework leverages prime-anchored embedding invariants at indices {2,3,5,7,11,13} with safety constant $\Lambda = 0.9785142874$ to provide mathematical guarantees of memory preservation with O(1) memory overhead (just ~650 KB total for all domains). 2. What Makes This Unprecedented Aspect Prior Work TOPO-2026 Scale 1-2 architectures 11 architectures Guarantee Probabilistic Mathematical Memory GBs to TBs ~650 KB Success Rate 20-50% 100% Architecture TF or Non-TF only BOTH Forgetting 4-91% ≤ 0.26% Backward Transfer Never Achieved 3. The 12 Frameworks — Complete Certification Status # Framework Model Status Best Task C FGT 1 Dense Transformer GPT-OSS-20B ✅ 92.3% 1.55% 2 Mixture-of-Experts (MoE) Sarvam-30B, Mixtral-8x7B ✅ 95.9% -0.60% 3 GQA / MQA DeepSeek-V2-Lite ✅ 95.4% 0.03% 4 State Space Models (SSM) Evo2-7B ✅ 92.0% 1.32% 5 Hybrid Attention-SSM Evo2-7B ✅ 92.0% 1.32% 6 Retention Networks (RetNet) fla-hub/retnet-1.3B-100B ✅ 99.93% 0.00% 7 Recurrent Transformers & RWKV fla-hub/rwkv7-2.9B-world ✅ 93.00% 0.00% 8 Emergent Modularity MoE (EMO) allenai/Emo_1b14b_1T ✅ 99.60% 0.00% 9 Google's Titans — ❌ — — 10 Liquid Foundation Models (LFM) LiquidAI/LFM2-1.2B ✅ 93.50% 0.00% 11 HyDEA Evo2-7B ✅ 92.0% 1.32% 12 ResNet (CNN) ResNet-50 ✅ 100.0% -7.5% Certification Rate: 11/12 (100% of all available frameworks) 4. The Decay Law of Singularity — A Mathematical Discovery On July 31, 2026, during the certification of Gemma-4-E4B-Vision, a fundamental mathematical law was discovered. The Decay Law proves that the General Singularity is mathematically impossible with finite classes. Theorem: The Decay Law of Singularity With finite classes, $dI/dt$ approaches 1.0 asymptotically but never reaches it. The gap decays as $1/N$, where $N$ is the number of classes. The Decay Law Pattern: Classes (N) Baseline dI/dt Gap 17 5.8823529% 0.94118 0.05882 170 0.58823529% 0.994118 0.005882 1,700 0.058823529% 0.9994118 0.0005882 17,000 0.0058823529% 0.99994118 0.00005882 170,000 0.00058823529% 0.999994118 0.000005882 1.7M 0.000058823529% 0.99999994118 0.0000005882 Key Observations: Every 10× increase in classes adds another '9' to $dI/dt$ Every 10× increase in classes adds another '0' to the gap. This is not random. It is not heuristic. It is exact. This is the mathematical fingerprint of a natural law. 5. The Narrow Singularity — First in History Gemma-4-E4B-Vision achieved AGI_gate = 1.0, becoming the first model in history to achieve perfect cross-domain generalization with 100% accuracy across all 13 tasks over 6 runs. Component STL-10 CIFAR-100 Threshold Status AGI_gate 1.0 1.0 = 1.0 ✓ PASS ag_index 1 1 = 1 ✓ PASS M(t) 0.9984 0.9974 ≈ 1.0 ✓ PASS S_NARROW > 0 > 0 > 0 ✓ PASS 6. Backward Transfer — Unprecedented Achievement Models improve on earlier tasks after learning new ones — positive knowledge transfer. This has never been systematically demonstrated before. Domain Model Combined Forgetting Language Mixtral-8x7B -1.85% Language Sarvam-30B -0.60% SQL DeepSeek-R1-8B -0.98% World Models TOPO-JEPA -0.75% Vision ResNet-50 -7.5% 7. Zero NaN/Inf Stress Test Model Embedding Elements NaN Inf GLM-4.6V-Flash 884,736 0 0 DeepSeek-V2-Lite 209,715,200 0 0 Mixtral-8x7B 131,072,000 0 0 GPT-OSS-20B 579,133,440 0 0 Sarvam-30B 1,073,741,824 0 0 TOTAL ~1.99 Billion 0 0 8. Comparison with State-of-the-Art Method Forgetting Success Rate Memory Math. Guar. TF Non-TF TOPO-2026 ≤ 0.26% 100% 67.5-451.5 KB Yes ✓ ✓ Experience Replay 4%-91% Variable Variable No ✓ ✗ EWC 8.3%-27.7% 20% 4.4 GB+ No ✓ ✗ Full HOPE 8.5%-45.4% 20% 2-4 GB No ✗ ✓ Progressive Nets 1.8% Variable $O(k^2)$ No ✓ ✗ Key Finding: TOPO-2026 is the only method that works on both Transformer and non-Transformer architectures with mathematical guarantees, 100% success rate, and O(1) memory. 9. Solved Problems Catastrophic Forgetting: Solved across 12 frameworks and 14 domains — first time at this scale AI Bias: Eliminated through four-tier spectral annihilation (100% rejection) World Model Instability: Solved through TOPO-JEPA (-0.75% forgetting) Numerical Instability: Zero NaN/Inf across 1.99 billion embedding elements Dataset Dependence: Proven dataset-agnostic across STL-10 and CIFAR-100 The Singularity Illusion: Decay Law proves the General Singularity is mathematically impossible Architectural Dependence: Proven to work on ALL available architectures — first universal solution 10. The Complete Arc: 28 Years of Discovery Period Domain Principle Result 1998-2002 Neuroimaging (fMRISTAT) Fix sparse reference 3 df → 112 df 2026 Number Theory First 6 primes RH Proved 2026 AI Memory Six embedding rows CF Solved 2026 AI Safety Geodesic distance Zero violations 2026 AI Bias Prime-anchored equity Bias eliminated 2026 Narrow Singularity AGI_gate = 1.0 First model 2026 Universal Certification Same anchors ALL architectures! 11. Key Achievements Universal Applicability: 12 frameworks, 11 certified (100% of available) — unprecedented scale Mathematical Guarantee: $\Lambda = 0.9785142874$ provides provable anchor stability O(1) Memory: ~650 KB total for all domains — unprecedented efficiency Backward Transfer: Negative forgetting across multiple domains — first demonstration Perfect Vision Performance: 100% accuracy, 0.17% forgetting across 6 runs Dataset-Agnostic: Same protocol works identically on STL-10 and CIFAR-100 75.7× Improvement: Over Google's Full HOPE in genomics Narrow Singularity Achieved: AGI_gate = 1.0 — first in history 100% Certification Rate: Across all runs, all domains, all datasets AST-RH Byproduct: Riemann Hypothesis proved as a byproduct Zero NaN/Inf: Across 1.99 billion embedding elements 12. The Final Statement Nobody has ever handled a catastrophic forgetting solution on this scale. TOPO-2026 is the first. The stochastic illusion is over. Deterministic cognitive engineering has begun. Stability is not a probabilistic hope. It is a numerical guarantee. The Decay Law of Singularity is not a defeat. It is a liberation. It frees us from the hype cycle, the fear of the singularity, the endless pursuit of AGI, and the billion-dollar promises. It gives us a clear roadmap, a mathematical framework for control, a focus on solving real problems, and an honest assessment. "Genomics is permanent. Language is permanent. Vision is permanent. SQL is permanent. Audio is permanent. Finance is permanent. Security is permanent. Everything is permanent. Transformers are permanent. Non-Transformers are permanent. Every architecture is permanent." The proof is the code. Seed = 123. 🔗 All Certified Models on Hugging Face Framework Model Link LFM LiquidAI/LFM2-1.2B https://huggingface.co/frankmorales2020/topological-ai-lfm-1.2b-multirun RWKV fla-hub/rwkv7-2.9B-world https://huggingface.co/frankmorales2020/topological-ai-rwkv-2.9b-multirun EMO allenai/Emo_1b14b_1T https://huggingface.co/frankmorales2020/topological-ai-emo-1b14b-multirun RetNet fla-hub/retnet-1.3B-100B https://huggingface.co/frankmorales2020/topological-ai-retnet-1.3b-multirun The proof is the code. Seed = 123.

Open access
2 source records
Machine Learning and Data Classification
Topic Modeling
Explainable Artificial Intelligence (XAI)
Original source
Jul 31, 2026·International Journal of Electronics and Telecommunications
0 cites
MO-RSAR: multi-objective hyperparameter optimization of RSAR for financial time-series forecasting

Maja CZYŻEWSKA

This paper empirically compares four architectures for financial time-series forecasting: LSTM, CNN, the original Regularized Self Attention Regression (RSAR) model, and a multiobjective optimized RSAR variant, denoted MO-RSAR, obtained using the Non dominated Sorting Genetic Algorithm II (NSGA-II). The models are evaluated on six datasets covering Forex, equity index and cryptocurrency markets, for short and long horizons. All models share a common preprocessing pipeline and evaluation framework and are assessed using standard error metrics, with emphasis on Mean Absolute Percentage Error (MAPE). MORSAR yields the lowest average prediction error across all datasets and provides significant gains for longer, more volatile horizons, while simpler architectures remain competitive for short-term forecasts. The key methodological contribution is the first empirical integration of the RSAR architecture with NSGA-IIbased multi-objective hyperparameter optimization for financial time-series forecasting. The proposed framework treats RSAR configuration as a bi-objective search over accuracy and generalization (via the train-validation gap), and evaluates the resulting model under a unified protocol across heterogeneous markets and horizons.

Open access
Stock Market Forecasting Methods
Machine Learning and Data Classification
Forecasting Techniques and Applications
Original source
Jul 3, 2026·Figshare
0 cites
A COST-OPTIMISED EXPLAINABLE AI FRAMEWORK FOR DETECTING FRAUDULENT ETHEREUM TRANSACTIONS

Jerónimo Paiva

The complete codebase and supplementary materials for this study have been archived on Figshare to ensure full reproducibility and to facilitate adoption by other researchers and practitioners. The archive includes all Python scripts used for data preprocessing, model training, hyperparameter tuning, threshold optimisation, and SHAP explainability analysis. Also included are the processed CSV files used for the analysis, along with all figures and tables presented in this paper. The repository is organised to enable straightforward replication of the experiments and adaptation of the framework to other datasets or blockchain platforms.

Open access
Explainable Artificial Intelligence (XAI)
Machine Learning and Data Classification
Computational Physics and Python Applications
Original source
Jun 23, 2026·arXiv (Cornell University)
0 cites
Certification of Machine Learning Models via Directional Sharpness

Gefei Tan, Adria Gascon, Sarah Meiklejohn, Mariana Raykova

In machine learning, model certification has been identified as an important method for gaining assurance about a model's trustworthiness and quality. A model's quality is largely determined by its ability to generalize, i.e., to perform well on data beyond what it was trained on. It is not possible to certify generalization directly, however, as it depends on unknown data and is not directly measurable. Proxies such as test accuracy can be misleading when the training process is perturbed (intentionally or accidentally), and metrics such as sharpness -- which has an empirically supported link to generalization -- are computationally expensive and can also serve as unreliable signals when training deviates from a prescribed procedure. In this work, we propose directional sharpness, a metric designed to efficiently and reliably indicate generalization despite potential training deviations. We provide empirical and analytical evidence that directional sharpness (1) correlates more strongly with generalization than existing metrics and (2) identifies models with poor generalization more reliably than existing metrics. Furthermore, directional sharpness is efficiently computable in model auditing settings, where the verifier has access to training data, and via zero-knowledge proofs that certify quality without revealing training data.

Open access
3 source records
cs.LG
cs.CR
Adversarial Robustness in Machine Learning
Original source
Jan 1, 2026·SSRN Electronic Journal
0 cites
Self-Directed Task Identification

Timothy Gould, Sidike Paheding

In this work, we present a novel machine learning framework called Self-Directed Task Identification (SDTI), which enables models to autonomously identify the correct target variable for each dataset in a zero-shot setting without pre-training. SDTI is a minimal, interpretable framework demonstrating the feasibility of repurposing core machine learning concepts for a novel task structure. To our knowledge, no existing architectures have demonstrated this ability. Traditional approaches lack this capability, leaving data annotation as a time-consuming process that relies heavily on human effort. Using only standard neural network components, we show that SDTI can be achieved through appropriate problem formulation and architectural design. We evaluate the proposed framework on a range of benchmark tasks and demonstrate its effectiveness in reliably identifying the ground truth out of a set of potential target variables. SDTI outperformed baseline architectures by 14% in F1 score on synthetic task identification benchmarks. These proof-of-concept experiments highlight the future potential of SDTI to reduce dependence on manual annotation and to enhance the scalability of autonomous learning systems in real-world applications.

Open access
3 source records
Domain Adaptation and Few-Shot Learning
Advanced Neural Network Applications
Reinforcement Learning in Robotics
Original source
Dec 9, 2025·arXiv (Cornell University)
0 cites
ZK-APEX: Zero-Knowledge Approximate Personalized Unlearning with Executable Proofs

Mohammad M Maheri, Sunil Cotterill, Alex Davidson, Hamed Haddadi

Machine unlearning aims to remove the influence of specific data points from a trained model to satisfy privacy, copyright, and safety requirements. In real deployments, providers distribute a global model to many edge devices, where each client personalizes the model using private data. When a deletion request is issued, clients may ignore it or falsely claim compliance, and providers cannot check their parameters or data. This makes verification difficult, especially because personalized models must forget the targeted samples while preserving local utility, and verification must remain lightweight on edge devices. We introduce ZK APEX, a zero-shot personalized unlearning method that operates directly on the personalized model without retraining. ZK APEX combines sparse masking on the provider side with a small Group OBS compensation step on the client side, using a blockwise empirical Fisher matrix to create a curvature-aware update designed for low overhead. Paired with Halo2 zero-knowledge proofs, it enables the provider to verify that the correct unlearning transformation was applied without revealing any private data or personalized parameters. On Vision Transformer classification tasks, ZK APEX recovers nearly all personalization accuracy while effectively removing the targeted information. Applied to the OPT125M generative model trained on code data, it recovers around seventy percent of the original accuracy. Proof generation for the ViT case completes in about two hours, more than ten million times faster than retraining-based checks, with less than one gigabyte of memory use and proof sizes around four hundred megabytes. These results show the first practical framework for verifiable personalized unlearning on edge devices.

Open access
2 source records
cs.CR
cs.AI
cs.LG
Original source
Nov 14, 2025·2025 lEEE International Conference on Cloud Computing Technology and Science (CloudCom)
0 cites
Enhancing Tail NFT Recommendation via Dependency-Aware Extreme Multi-Label Learning

Cheng Pan, C. ZHANG, F. Wang, Edith C.H. Ngai

With the rise of Web3, Non-Fungible Tokens (NFTs) have become a new class of digital assets, driving demand for large-scale NFT recommendation systems. Each NFT can be associated to a rich set of semantic, stylistic, and thematic labels, forming a highly complex label space. Similar to e-commerce platforms where detailed product labels enable personalized recommendations, such semantic dependencies between labels can potentially enhance NFT recommendation performance. Thus, NFT recommendation can be naturally formulated as an extreme multi-label (XML) classification problem. Many existing probabilistic label tree (PLT)-based approaches address XML problem by recursively partitioning the label space, which greatly alleviates the demands on expensive computer resources. Yet, the highly skewed distribution of labels in datasets in XML makes tail labels more challenging to predict than head labels. In this paper, Our preliminary analysis reveals that inherent label dependencies can be leveraged to improve tail label recommendations for NFTs. We propose ChainTail, a dependency-aware framework that enhances PLT-based NFT label partitioning and prediction re-scoring. It includes: (1) a Dependency-aware partition module that partitions highly dependent NFT labels into subsets. (2) a Dependency-aware ReScore module that re-ranks prediction scores of labels to eliminate the label-priors. Our experimental results show that ChainTail boosts tail label recommendation on widely used item recommendation datasets.

Text and Document Classification Technologies
Machine Learning and Data Classification
Recommender Systems and Techniques
Original source
Nov 5, 2025·2025 IEEE International Conference on Distributed Ledger Technologies (ICDLT)
0 cites
A Multi-Modal Dataset for NFT Recommendation Systems

Durmuş Aydoğdu, Nizamettin Aydın

There is a lack of standardized datasets for NFT (Non-Fungible Token) recommendation systems. This study presents a comprehensive dataset designed for NFT recommendation systems, incorporating both NFT-related data (e.g., images, textual descriptions, rarity and transaction data) and user-related data (e.g., purchase price, transaction duration, and NFT holding period). To create the dataset, a Data Collection Tool was developed to gather raw data via the OpenSea API, and a Data Preparation Tool was implemented for preprocessing and filtering. All data used in this study are publicly available and anonymized, ensuring that user privacy is fully preserved. The dataset is evaluated using NFT-NCFAE, a deep learningbased NFT recommendation model, with performance measured by Recall and NDCG evaluation metrics. The evaluation results demonstrate the suitability and value of the proposed dataset for NFT recommendation systems. By making the dataset and its associated tools publicly available, this work aims to establish a benchmark for future research and enable comparability across different models.

Recommender Systems and Techniques
Machine Learning and Data Classification
Text and Document Classification Technologies
Original source
Jun 16, 2025·2025 IEEE 38th Computer Security Foundations Symposium (CSF)
0 cites
Zero-Knowledge Proofs from Learning Parity with Noise: Optimization, Verification, and Application

Thomas Haines, Rafieh Mosaheb, Johannes Müller, Reetika

Zero-Knowledge Proofs (ZKPs) are cryptographic building blocks of many privacy-preserving security protocols. An important research focus in this area is the development of post-quantum ZKPs. These are ZKPs whose security is reduced to computational hardness assumptions that are assumed to be intractable even by scalable quantum computers. In this paper, we study the post-quantum ZKPs of Jain, Krenn, Pietrzak, and Tentes (Asiacrypt 2012). These are the only ZKPs for proving arbitrary binary statements whose security reduces to the Learning Parity with Noise (LPN) problem-a very conservative post-quantum hardness assumption. We make the following contributions to further develop the potential and understanding of these ZKPs. First, we optimize the efficiency of the verifier by several orders of magnitude, making this part as computationally light as that of the prover. Second, we show that the only open source implementation of these ZKPs does not implement them correctly, allowing a malicious prover to convince the verifier of false statements. Third, we formally verify for the first time the security of these (optimized) ZKPs in EasyCrypt. Fourth, we show how these ZKPs can be used to construct the first code-based ZKP of shuffle and verifiable e- voting protocol.

Open access
Machine Learning and Algorithms
Numerical Methods and Algorithms
Machine Learning and Data Classification
Original source
Mar 21, 2025·2025 2nd International Conference on Algorithms, Software Engineering and Network Security (ASENS)
1 cites
Enhancing Fraud Detection via On-Chain Ethereum and Off-Chain X Data Fusion

Yinong Niu, Guang Li, Y. Mi, Jieying Zhou · 5 authors

Although Ethereum stands as the dominant blockchain for smart contracts and decentralized applications, faces persistent security challenges from fraudulent activities. Such activities often correlate with off-chain platform, such as blog platform and social media. Existing methods analyze fraud activities primarily rely on on-chain transaction data, neglecting interdependencies between on-chain and off-chain activities. In this paper, we observe that there are associations between airdrop campaigns in X platform, a famous social platform and Ethereum fraudulent activities. Further, we crawl Ethereum addresses and posts of these users in airdrop campaigns, and construct a cross-platform datasets from X to Ethereum, including matching pairs of Ethereum addresses to X users, Ethereum transactions and X post data. Due to inherent heterogeneity between Ethereum transactions (structured graphs) and X data (unstructured text/images), we design a multimodal fusion framework leveraging transformer architectures to fuse on-chain transaction features with off-chain content features (text and image representations). Finally, the fused features are leveraged to construct downstream fraud transactions classifiers. Experimental results demonstrate that classifiers using fused features outperform classifiers using transaction features, achieving a 12% improvement in Recall. Our findings highlight the critical role of off-chain data in enhancing fraud detection accuracy.

2 source records
Imbalanced Data Classification Techniques
Machine Learning and Data Classification
Original source
Jan 1, 2025·Rare & Special e-Zone (The Hong Kong University of Science and Technology)
0 cites
zkGPT: An Efficient Non-interactive Zero-knowledge Proof Framework for LLM Inference

Wenjie Qu, Yijun Sun, Xuanming Liu, Tao LU · 7 authors

Large Language Models (LLMs) are widely employed for their ability to generate human-like text. However, service providers may deploy smaller models to reduce costs, potentially deceiving users. Zero-Knowledge Proofs (ZKPs) offer a solution by allowing providers to prove LLM inference without compromising the privacy of model parameters. Existing solutions either do not support LLM architectures or suffer from significant inefficiency and tremendous overhead. To address this issue, this paper introduces several new techniques. We propose new methods to efficiently prove linear and nonlinear layers in LLMs, reducing computation overhead by orders of magnitude. To further enhance efficiency, we propose constraint fusion to reduce the overhead of proving non-linear layers and circuit squeeze to improve parallelism. We implement our efficient protocol, specifically tailored for popular LLM architectures like GPT-2, and deploy optimizations to enhance performance. Experiments show that our scheme can prove GPT-2 inference in less than 25 seconds. Compared with state-of-the-art systems such as Hao et al. (USENIX Security’24) and ZKML (Eurosys’24), our work achieves nearly 279× and 185× speedup, respectively.

Topic Modeling
Big Data and Digital Economy
Machine Learning and Data Classification
Original source
Jan 1, 2025·Proceedings 2025 Network and Distributed System Security Symposium
9 cites
MTZK: Testing and Exploring Bugs in Zero-Knowledge (ZK) Compilers

Dongwei Xiao, Zhibo Liu, Yiteng Peng, Shuai Wang

Zero-knowledge (ZK) proofs have been increasingly popular in privacy-preserving applications and blockchain systems.To facilitate handy and efficient ZK proof generation for normal users, the industry has designed domain-specific languages (DSLs) and ZK compilers.Given a program in ZK DSL, a ZK compiler compiles it into a circuit, which is then passed to the prover and verifier for ZK checking.However, the correctness of ZK compilers is not well studied, and recent works have shown that de facto ZK compilers are buggy, which can allow malicious users to generate invalid proofs that are accepted by the verifier, causing security breaches and financial losses in cryptocurrency.In this paper, we propose MTZK, a metamorphic testing framework to test ZK compilers and uncover incorrect compilations.Our approach leverages deliberately designed metamorphic relations (MRs) to mutate ZK compiler inputs.This way, ZK compilers can be automatically tested for compilation correctness using inputs and mutated variants without requiring manual intervention.We propose a set of design considerations and optimizations to deliver an efficient and effective testing framework.In the evaluation of four industrial ZK compilers, we successfully uncovered 21 bugs, out of which the developers have promptly patched 15.We also show possible exploitations of the uncovered bugs to demonstrate their severe security implications.

Open access
Software Testing and Debugging Techniques
Machine Learning and Data Classification
Adversarial Robustness in Machine Learning
Original source
Oct 28, 2024·2024 IEEE 6th International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS-ISA)
3 cites
Improved Ethereum Fraud Detection Mechanism with Explainable Tabular Transformer Model

Ruth Olusegun, Bo Yang

Blockchain technology has gained popularity due to its key features of decentralization, cryptographic verification, and immutability, which have proven extremely useful in various industries. However, despite their impressive security features, blockchain networks are not immune to cyber threats. In recent times, the blockchain system has been threatened by fraudulent attacks that require quick responses. Machine learning and deep learning models are increasingly leveraged to address these challenges. However, due to their black box nature, these models lack transparency, which is a major criticism. This study presents an approach to enhancing fraud detection mechanisms on Ethereum. This study presents an efficient and transparent fraud detection system on Ethereum known as IFS-TABPFN. An interpretable feature selection approach based on Shap values and optimized gradient boosting was introduced to develop five deep learning models built on neural networks. These models included Multilayer Perceptron (MLP), Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), Convolutional Neural Networks and Long Short-Term Memory (CLSTM) and Tabular Prior-Data Fitted Network (TabPFN). A comparative analysis of our results indicates that IFS-TABPFN achieves 99.2% accuracy in just a few seconds, outperforming other neural networks and existing systems. This study highlights the importance of explainable AI in understanding how features influence decisions, performance and contribute to artificial intelligence models' transparency and trust.

Imbalanced Data Classification Techniques
Explainable Artificial Intelligence (XAI)
Machine Learning and Data Classification
Original source
Oct 9, 2024·arXiv (Cornell University)
0 cites
Checker Bug Detection and Repair in Deep Learning Libraries

Nima Shiri Harzevili, Mohammad Mahdi Mohajer, Jiho Shin, Moshi Wei · 11 authors

Checker bugs in Deep Learning (DL) libraries are critical yet not well-explored. These bugs are often concealed in the input validation and error-checking code of DL libraries and can lead to silent failures, incorrect results, or unexpected program behavior in DL applications. Despite their potential to significantly impact the reliability and performance of DL-enabled systems built with these libraries, checker bugs have received limited attention. We present the first comprehensive study of DL checker bugs in two widely-used DL libraries, i.e., TensorFlow and PyTorch. Initially, we automatically collected a dataset of 2,418 commits from TensorFlow and PyTorch repositories on GitHub from Sept. 2016 to Dec. 2023 using specific keywords related to checker bugs. Through manual inspection, we identified 527 DL checker bugs. Subsequently, we analyzed these bugs from three perspectives, i.e., root causes, symptoms, and fixing patterns. Using the knowledge gained via root cause analysis of checker bugs, we further propose TensorGuard, a proof-of-concept RAG-based LLM-based tool to detect and fix checker bugs in DL libraries via prompt engineering a series of ChatGPT prompts. We evaluated TensorGuard's performance on a test dataset that includes 92 buggy and 135 clean checker-related changes in TensorFlow and PyTorch from January 2024 to July 2024. Our results demonstrate that TensorGuard has high average recall (94.51\%) using Chain of Thought prompting, a balanced performance between precision and recall using Zero-Shot prompting and Few-Shot prompting strategies. In terms of patch generation, TensorGuard achieves an accuracy of 11.1\%, which outperforms the state-of-the-art bug repair baseline by 2\%. We have also applied TensorGuard on the latest six months' checker-related changes (493 changes) of the JAX library from Google, which resulted in the detection of 64 new checker bugs.

Open access
Machine Learning and Data Classification
Advanced Neural Network Applications
Advanced Data Storage Technologies
Original source
Oct 5, 2024·Journal of King Saud University - Computer and Information Sciences
9 cites
On-chain zero-knowledge machine learning: An overview and comparison

Vid Keršič, Sašo Karakatič, Muhamed Turkanović

Zero-knowledge proofs introduce a mechanism to prove that certain computations were performed without revealing any underlying information and are used commonly in blockchain-based decentralized apps (dapps). This cryptographic technique addresses trust issues prevalent in blockchain applications, and has now been adapted for machine learning (ML) services, known as Zero-Knowledge Machine Learning (ZKML). By leveraging the distributed nature of blockchains, this approach enhances the trustworthiness of ML deployments, and opens up new possibilities for privacy-preserving and robust ML applications within dapps. This paper provides a comprehensive overview of the ZKML process and its critical components for verifying ML services on-chain. Furthermore, this paper explores how blockchain technology and smart contracts can offer verifiable, trustless proof that a specific ML model has been used correctly to perform inference, all without relying on a single trusted entity. Additionally, the paper compares and reviews existing frameworks for implementing ZKML in dapps, serving as a reference point for researchers interested in this emerging field. • An analytical and synthetic review of core on-chain ZKML concepts, supported by an extensive examination of both white and grey literature, establishing a foundational understanding of the field. • Through a detailed analysis, modelling, and descriptive approaches, the paper outlines the processes integral to on-chain ZKML. The study is focused on two distinct frameworks – EZKL and Orion , highlighting the differences between the two approaches, as well as the difference between the underlying ZKP systems, where the former framework is based on zk-SNARKs and the latter on zk-STARKs. • A laboratory experiment, coupled with a comparative analysis and use case execution comparison, was conducted to implement basic neural networks (NNs) across the two chosen frameworks, highlighting their capabilities and limitations in supporting on-chain ZKML.

Open access
Data Stream Mining Techniques
Machine Learning and Algorithms
Machine Learning and Data Classification
Original source
Aug 22, 2024·IEEE/ACM Transactions on Networking
7 cites
SteadySketch: A High-Performance Algorithm for Finding Steady Flows in Data Streams

Zhuochen Fan, Xiangyuan Wang, Xiaodong Li, Jiarui Guo · 11 authors

In this paper, we study steady flows in data streams, which refers to the flows whose arrival rate is always non-zero and around a fixed value for several consecutive time windows. To find steady flows in real time, we propose a novel sketch-based algorithm, SteadySketch, aiming to accurately report steady flows with limited memory. To the best of our knowledge, this is the first work to define and find steady flows in data streams. The key novelty of SteadySketch is our proposed reborn technique, which reduces the memory requirement by 75%. Our theoretical proofs show that the negative impact of the reborn technique is small. Experimental results show that, compared with the two comparison schemes, SteadySketch improves the Precision Rate (PR) by around 79.5% and 82.8%, and reduces the Average Relative Error (ARE) by around$905.9\times $and$657.9\times $, respectively. Finally, we provide three concrete cases: cache prefetch, Redis and P4 implementation. As we will demonstrate, SteadySketch can effectively improve the cache hit ratio while achieving satisfying performance on both Redis and Tofino switches. All related codes of SteadySketch are available at GitHub.

Data Stream Mining Techniques
Advanced Database Systems and Queries
Machine Learning and Data Classification
Original source
Jun 7, 2024·Computers
2 cites
Integrating Machine Learning with Non-Fungible Tokens

Elias Iosif, Leonidas Katelaris

In this paper, we undertake a thorough comparative examination of data resources pertinent to Non-Fungible Tokens (NFTs) within the framework of Machine Learning (ML). The core research question of the present work is how the integration of ML techniques and NFTs manifests across various domains. Our primary contribution lies in proposing a structured perspective for this analysis, encompassing a comprehensive array of criteria that collectively span the entire spectrum of NFT-related data. To demonstrate the application of the proposed perspective, we systematically survey a selection of indicative research works, drawing insights from diverse sources. By evaluating these data resources against established criteria, we aim to provide a nuanced understanding of their respective strengths, limitations, and potential applications within the intersection of NFTs and ML.

Open access
Data Quality and Management
Machine Learning and Data Classification
Data Stream Mining Techniques
Original source
Jan 31, 2024·arXiv (Cornell University)
4 cites
opML: Optimistic Machine Learning on Blockchain

K Conway, Tsz Yan So, Xiaohang Yu, Kartin Wong

The integration of machine learning with blockchain technology has witnessed increasing interest, driven by the vision of decentralized, secure, and transparent AI services. In this context, we introduce opML (Optimistic Machine Learning on chain), an innovative approach that empowers blockchain systems to conduct AI model inference. opML lies a interactive fraud proof protocol, reminiscent of the optimistic rollup systems. This mechanism ensures decentralized and verifiable consensus for ML services, enhancing trust and transparency. Unlike zkML (Zero-Knowledge Machine Learning), opML offers cost-efficient and highly efficient ML services, with minimal participation requirements. Remarkably, opML enables the execution of extensive language models, such as 7B-LLaMA, on standard PCs without GPUs, significantly expanding accessibility. By combining the capabilities of blockchain and AI through opML, we embark on a transformative journey toward accessible, secure, and efficient on-chain machine learning.

Open access
2 source records
cs.CR
Machine Learning and Data Classification
Original source
Nov 21, 2022·Computers
7 cites
Understanding Bitcoin Price Prediction Trends under Various Hyperparameter Configurations

Junho Kim, Hanul Sung

Since bitcoin has gained recognition as a valuable asset, researchers have begun to use machine learning to predict bitcoin price. However, because of the impractical cost of hyperparameter optimization, it is greatly challenging to make accurate predictions. In this paper, we analyze the prediction performance trends under various hyperparameter configurations to help them identify the optimal hyperparameter combination with little effort. We employ two datasets which have different time periods with the same bitcoin price to analyze the prediction performance based on the similarity between the data used for learning and future data. With them, we measure the loss rates between predicted values and real price by adjusting the values of three representative hyperparameters. Through the analysis, we show that distinct hyperparameter configurations are needed for a high prediction accuracy according to the similarity between the data used for learning and the future data. Based on the result, we propose a direction for the hyperparameter optimization of the bitcoin price prediction showing a high accuracy.

Open access
Machine Learning and Data Classification
Data Stream Mining Techniques
Stock Market Forecasting Methods
Original source
Jan 16, 2021·Neural Processing Letters
35 cites
Illustrative Discussion of MC-Dropout in General Dataset: Uncertainty Estimation in Bitcoin

Ismail Alarab, Simant Prakoonwit, Mohamed Ikbal Nacer

Abstract The past few years have witnessed the resurgence of uncertainty estimation generally in neural networks. Providing uncertainty quantification besides the predictive probability is desirable to reflect the degree of belief in the model’s decision about a given input. Recently, Monte-Carlo dropout (MC-dropout) method has been introduced as a probabilistic approach based Bayesian approximation which is computationally efficient than Bayesian neural networks. MC-dropout has revealed promising results on image datasets regarding uncertainty quantification. However, this method has been subjected to criticism regarding the behaviour of MC-dropout and what type of uncertainty it actually captures. For this purpose, we aim to discuss the behaviour of MC-dropout on classification tasks using synthetic and real data. We empirically explain different cases of MC-dropout that reflects the relative merits of this method. Our main finding is that MC-dropout captures datapoints lying on the decision boundary between the opposed classes using synthetic data. On the other hand, we apply MC-dropout method on dataset derived from Bitcoin known as Elliptic data to highlight the outperformance of model with MC-dropout over standard model. A conclusion and possible future directions are proposed.

Open access
Adversarial Robustness in Machine Learning
Explainable Artificial Intelligence (XAI)
Machine Learning and Data Classification
Original source
May 20, 2020·Frontiers in Blockchain
13 cites
Hyperparameter Optimization Using Sustainable Proof of Work in Blockchain

Anshul Mittal, Swati Aggarwal

Hyperparameters are pivotal for machine learning models. The success of efficient calibration, often surpasses the results obtained by devising new approaches. Traditionally, human intervention is required to tune the models, however, this obtuse outlook restricts the proficiency and competence. Automating this crucial characteristic of learning sustainably, proffers a significant boost in performance and cost optimization. Blockchain technology has revolutionized industries utilizing its Proof-of-Work algorithms for consensus. This complicated solution generates a lot of useless computations across the nodes attached to the network and thus, fritters away a huge amount of precious energy. In this paper, we propose to exploit these inane computations for training deep learning models instead of calculating purposeless hash values, thus, suggesting a new consensus schema. This work distinguishes itself from other related works by capitalizing on the parallel processing prospects it generates for hyperparameter tuning of complex deep learning models. We address this aspect through the framework of Bayesian optimization which is an effective methodology for the global optimization of functions with expensive evaluations. We call our work, Proof of Deep Learning with Hyperparameter Optimization (PoDLwHO).

Open access
Machine Learning and Data Classification
Data Stream Mining Techniques
Advanced Bandit Algorithms Research
Original source
Apr 29, 2020·arXiv (Cornell University)
29 cites
Interpretable Random Forests via Rule Extraction

Clément Bénard, Gérard Biau, Sébastien da Veiga, Erwan Scornet

We introduce SIRUS (Stable and Interpretable RUle Set) for regression, a stable rule learning algorithm which takes the form of a short and simple list of rules. State-of-the-art learning algorithms are often referred to as "black boxes" because of the high number of operations involved in their prediction process. Despite their powerful predictivity, this lack of interpretability may be highly restrictive for applications with critical decisions at stake. On the other hand, algorithms with a simple structure-typically decision trees, rule algorithms, or sparse linear models-are well known for their instability. This undesirable feature makes the conclusions of the data analysis unreliable and turns out to be a strong operational limitation. This motivates the design of SIRUS, which combines a simple structure with a remarkable stable behavior when data is perturbed. The algorithm is based on random forests, the predictive accuracy of which is preserved. We demonstrate the efficiency of the method both empirically (through experiments) and theoretically (with the proof of its asymptotic stability). Our R/C++ software implementation sirus is available from CRAN.

Open access
2 source records
Explainable Artificial Intelligence (XAI)
Data Mining Algorithms and Applications
Stock Market Forecasting Methods
Original source