Rohan Goyal, Venkatesan Guruswami, Yihang Sun, Mary Wootters
Proximity gaps are a property of error correcting codes that arise in the study of Interactive Oracle Proofs (IOPs) and Succinct Non-interactive Arguments of Zero Knowledge (SNARKs). Recent work of Goyal and Guruswami has established near-optimal proximity gaps for many families of codes, including subspace design codes, as well as random ensembles like random linear codes, Reed-Solomon codes with random evaluation points, and Gallager's ensemble of LDPC codes (Goyal & Guruswami, 2025). However, the parameters for these latter randomized ensembles are worse than the parameters for subspace design codes, and degrade as the degree ell increases. In this work, we obtain improved proximity gaps for random ensembles of codes, including random linear codes, Reed-Solomon codes with random evaluation points, and Gallager's ensemble. Quantitatively, our results for these random ensembles match the results that Goyal and Guruswami attained for subspace design codes. In fact, our techniques are a black-box transference from subspace design codes: any progress on subspace design codes will automatically lead to analogous progress for these random ensembles. To obtain our results, we extend the Local Coordinate-wise Linear (LCL) property framework developed by Levi, Mosheiff, and Shagrithaya and by Brakensiek, Chen, Dhar, and Zhang to a \textit{row-span constrained} version (Levi, Mosheiff & Shagrithaya, 2025; Brakensiek, Chen, Dhar & Zhang, 2025). This allows us to cast \textit{curve-decodability} -- a property that implies proximity gaps -- directly as a row-span constrained LCL property, and make use of that machinery. In contrast, because curve-decodability is not obviously a vanilla LCL property, prior work had worked with a proxy property instead, leading to the aforementioned parameter losses.
Context The exponential evolution and widespread integration of Artificial Intelligence (AI) and Machine Learning (ML) systems have fundamentally transformed industries, establishing AI as a central component in decision-making processes, task automation, and the optimization of complex operational pipelines. From healthcare diagnostics to financial forecasting and increasingly across critical cybersecurity infrastructure such as intrusion detection systems and malware classifiers, AI models are being deployed in environments where the correctness and authenticity of their outputs carry direct operational and safety consequences. Nevertheless, as the deployment of AI systems becomes widespread, the conditions under which these models are trained have evolved in a direction where the security landscape of them radically changes. The traaditional assumption of a centralized, fully controlled training environment, where a single trusted entity acquires data, trains the model, and deploys it, no longer reflects the reality of modern machine learning practice. The frequent use of remote sensing, federated learning and/or outsourced machine learning has introduced architectures where the entity that acquires the data, the entity that trains the model and the entity that ultimately relies on the model's output are three distinct and mutually distrusting parties. In a remote sensing scenario, sensors owned by a data provider transmit raw measurements to a training node that may be geographically or administratively distant. In a federated learning scenario, multiple decentralized devices train local models on their private data and submit the results to a central aggregator. In an outsourced learning scenario, a resource-constrained model sponsor delegates the training computation entirely to a third-party cloud provider. In all three cases, the common factor is the same: the model sponsor, the entity that is ultimately responsible for and dependent on the trained model, that does not control the data acquisition process, does not observe the training execution and has no native mechanism to verify that the model they receive is the result of the computation they requested, performed on the data they provided. This separation of control is the main focus addressed by this dissertation. It is not merely a theoretical concern: the literature has documented a wide range of attacks that exploit precisely this gap. When a malicious trainer substitutes data, alters labels, ignores some dataset's subsets or modifies model parameters, the resulting model may appear functionally correct on standard evaluation metrics while being systematically compromised for specific classes of input, an attack vector particularly dangerous in cybersecurity applications where a model that has been quietly trained to misclassify a specific type of malicious traffic provides no observable anomaly until the attack it was designed to hide occurs. Problem and Motivation The main motivation of this dissertation can be addressed as follows. Given a sensor, that produces a set of data points in a given time frame, or a dataset owned by a data provider and a model computed by a model trainer from that data, the model sponsor wants to ensure that the trained model is the result of executing a known training process over the complete and authenticated dataset $D_t$. That is, all data points in $D_t$ and only those data points were used as the training set. No modifications were made to those points or their labels and the obtained model is indeed the result obtained from the execution of the agreed training algorithm. This guarantee cannot be provided by standard Machine Learning procedures, like accuracy, precision or F1-score. A malicious trainer can submit a model that passes all the standard evaluation metrics on benign inputs while maintaining a targeted misclassification on a specific attack pattern. The only way to close this gap is to make the training process itself verifiable by requiring the trainer to produce and submit a cryptographic proof that is mathematically impossible to forge without having correctly executed the agreed computation on the authenticated data. This verification challenge comes together with a second problem, the \emph{model integrity gap} that exists between a trained model and its deployed representation. Even if the training process was all validated, the model must subsequently be transpiled and deployed into a certain non-ML format. In the context of this dissertation, this gap is particularly sensitive, the Python model trained by the data scientist must be translated into a ZoKrates arithmetic circuit for zero-knowledge proof generation, a process that involves converting continuous floating-point decision boundaries into discrete integer arithmetic. If this translation introduces a small inversion in a comparison operator or a shifted threshold values, the deployed circuit will produce systematically different predictions from the intended model and standard testing may not surface the discrepancy. The literature has proposed cryptographic solutions to the verifiable training but has largely left the second problem unaddressed. The foundational work by Keshavarzkalhori et al. demonstrated that it is possible to construct a pipeline combining hash chains, digital signatures and zero-knowledge proofs to verify that a simulated Naive Bayes classifier was trained on authenticated sensor data. Their implementation, built on the ZoKrates toolset, provided a proof-of-concept that the building blocks exist for end-to-end training verification. However, scaling this approach from a simple probabilistic classifier to a more complex, non-linear ensemble model, in this specific case, a Random Forest, introduces severe architectural bottlenecks that their work explicitly identified as open problems: the computational overhead of bitwise hashing inside arithmetic circuits, the floating-point to integer translation problem and the absence of any mechanism to verify that the transpilation of the model into the circuit was performed faithfully. This dissertation directly addresses these open problems. It proposes, implements and evaluates an end-to-end verifiable machine learning architecture for Random Forest classifiers that provides mathematical guarantees over three distinct integrity boundaries: the origin of the training data, the correctness of the training computation and the fidelity of the model's translation into a verifiable circuit. The framework is evaluated on both a simulated sensor dataset used by Keshavarzkalhori et al. and the CICIDS2017 network intrusion detection benchmark, the real-world cybersecurity dataset used by the most directly comparable prior work, demonstrating that the proposed integrity guarantees are achievable at practical computational cost for cybersecurity-relevant workloads. Research Questions The main objective of this thesis was to build a framework capable of protecting the overall AI Models from data and model poisoning attacks. In alignment with the goal, four research questions were set: Research Question 01: What state-of-the-art mechanisms exist to verify the integrity of AI models across the training pipeline? Research Question 02: What threats exist against AI models integrity? Research Question 03: What computational overhead do integrity verification mechanisms introduce across the AI modeling pipeline and how does this overhead scale with model complexity?
The rapid evolution of financial technology has transformed the global financial landscape, creating opportunities for innovation, inclusion, and efficiency while introducing systemic risks, regulatory uncertainties, and challenges to financial stability. This study presents a bibliometric review of global research trends at the intersection of financial technology and financial stability from 2000 to 2025, mapping the intellectual structure, identifying emerging themes, and highlighting influential contributions. Using Scopus data, the analysis examines 339 peer-reviewed documents across 242 sources. Bibliometric techniques were applied through VOSviewer, Bibliometrix (R), and Biblioshiny to evaluate publication trends, influential authors, thematic clusters, co-authorship networks, and keyword co-occurrences. The results show an average annual growth rate of 21.46 percent, with a marked increase in publications after 2017 coinciding with the mainstream adoption of digital finance and heightened policy focus on financial resilience. Findings indicate that financial technology promotes financial inclusion, banking efficiency, and economic empowerment, yet also introduces cybersecurity threats, regulatory gaps, and systemic vulnerabilities, particularly in emerging markets. Dominant themes include blockchain, digital payments, financial literacy, and central bank digital currencies, with decentralized finance and artificial intelligence emerging as fast-growing areas of scholarly interest. Geographically, China leads in publication volume, while the United Kingdom and the United States dominate in scholarly influence. This review provides a strategic roadmap for researchers and policymakers to navigate the evolving financial technology landscape and emphasizes the need for future research to integrate ethical governance, artificial intelligence risk management, and inclusive financial innovation frameworks.
Modular exponentiation is among the most demanding computational operations in cryptographic systems. Effective computation of modular exponentiation is most beneficial for public-key cryptography. The computational complexity and the growing number of bits of the key size, as required by increasingly stringent security demands in the RSA, the Diffie–Hellman key exchange and the Zero-Knowledge Proof (ZKP) protocols, have become a top research priority in terms of algorithmic efficiency. This study proposes a novel triple modular exponentiation algorithm based on the Improved Common-Multiplicand-Multiplication (ICMM) framework. The exact complexity formula was obtained through systematic probabilistic analysis of eight mutually exclusive bit-level states. The efficiency of modular exponentiation is primarily determined by the number of modular multiplications and exponentiation squares required. It is observed that improved common-multiplicand multiplication efficiently minimizes the computational complexity of the triple modular exponentiation by reducing the number of modular multiplications. The overall computational complexity of triple modular exponentiation is 1.875j, where j is the bit length of the exponent. This represents a reduction of approximately 16.7% in total multiplication count relative to double modular exponentiation, corresponding to a 44.4% reduction on a per-exponent basis, and a reduction of 58.3% relative to three independent binary exponentiations. This study concludes that the proposed decomposition reduces the average-case computational complexity of triple modular exponentiation to 1.875j modular multiplications for a j-bit exponent. The proposed triple modular exponentiation algorithm is shown to have lower number of multiplications per bit length of exponent as compared to double modular exponentiation. This result demonstrates the potential of proposed algorithm to reduce the computational cost of triple modular exponentiation in cryptographic protocols where it is a recurring operation, such as interactive ZKP identification schemes.
Ravindran Kandasamy, Chandan Chavadi, H. Chittoo, Nidhi Shukla
Online commerce, despite its infinite development possibilities, now raises the specter of global security. A huge amount of personal data is at risk from cyberattacks, such as hacking and identity theft, that harm companies and consumers alike. The traditional way of keeping everything in one place leads to unauthorized access and manipulation, thus requiring stronger security measures. The same decentralized, unbreakable encryption and immutable record keeping that give these barter platforms strong protection against fraud are also features of distributed ledger technology. Decentralization removed control from one single source, making it less likely that there will be any tampering and deception will become slim. Blockchain networks featuring “smart contracts” that make the terms of a deal transparent and enforce contracts without the need for go-betweens. This chapter provides an analysis of how the blockchain can enhance e-privacy in e-commerce, with a focus on the foundations and attributes of blockchain to overcome current threats. As the technology becomes widespread, real cases are proving to revolutionize data security. New Use Cases And Research Using Distributed Ledgers For Enhanced Security.
Paper 114BN addresses the mechanism gap left open by Paper 114BM. Paper 114BM showed that a frozen contact-deficit law could organize two-unit string-tension ratios in pure-gauge lattice theory while surviving named controls and a no-fit firewall. Paper 114BN asks whether that law can be explained from a primitive Holosphere-QCD contact-sharing ledger rather than treated only as a successful bridge expression. The paper models a two-unit flux object as two fundamental support lanes sharing a finite contact registry. In the large-color limit, the two lanes behave like independent fundamental strings. At finite color number, the two lanes lose some independence through two leading burdens: angular sharing and local contact overlap. The angular-sharing channel is interpreted as a one-lane finite-registry burden distributed over a half-turn exchange. The contact-overlap channel is interpreted as a two-lane coincidence burden. Together, these two channels reproduce the frozen Paper 114BM law without adding fitted coefficients or correction terms. The result is a conditional primitive mechanism derivation. It strengthens the Paper 114BM bridge by giving a compact reason for the two-channel structure, the angular normalization, and the contact-overlap scaling. It does not claim to derive continuum Yang-Mills theory, prove the mass gap, derive all string tensions, or establish full physical QCD. The main remaining task is to realize the same contact-sharing ledger on an explicit support graph and then test whether any higher-flux extension follows before target scoring.
Open access
2 source records
Quantum Chromodynamics and Particle Interactions
High-Energy Particle Collisions Research
Particle physics theoretical and experimental studies
The Al-Rakhawy Document for Digital Sovereignty (EPSA) presents a complete engineering blueprint for encrypted machine learning. It integrates Federated Learning, Zero-Knowledge Proofs, and Smart Contracts across five layers. Key innovations include Pedersen Commitments for lightweight edge processing and the Al-Rakhawy Equation, which calculates fair rewards based on marginal impact. This system ensures absolute data privacy, breaks central monopolies, and provides users with immediate, mathematically guaranteed economic returns.
Pawan Kumar Sanjaya, Christina Giannoula, Valdy Oktavian, Mehdi Saeedi · 7 authors
Zero-knowledge machine learning (zkML) enables a server to perform verifiable inference while keeping model parameters private from the client. However, existing zkML systems incur prohibitive proof-generation costs. We observe that proof generation exhibits limited parallelism; that is, prover time does not decrease significantly as the number of threads increases. This limitation is because existing systems rely on monolithic proof computation, constructing a single proof for the entire machine learning model. We introduce zkComposer, a modular proof-construction framework that unlocks an additional dimension of parallelism, in addition to the parallelism in existing proof kernels. zkComposer decomposes the zkML proof of correct inference into independent sub-proofs, each covering a subset of the computation for inference e.g., each independent sub-proof can cover a subset of contiguous layers in the ML model. Adjacent sub-proofs are cryptographically linked through shared commitments to the activations from the boundary layer. zkComposer provides the same guarantees as the monolithic proof without requiring additional linking proofs or changes to the underlying cryptographic primitives. We implement zkComposer and evaluate it on three CNNs and GPT-2. We show that, on CNN workloads, zkComposer reduces prover time and response time by up to 3.25x relative to zkCNN [1]. On GPT-2, zkComposer reduces these times by up to 4.83x relative to zkGPT [2], when partitioning along the model layers. When partitioning across both model layers and input sequences in GPT-2, we show that zkComposer reduces prover time and response time by up to 6.84x relative to zkGPT [2].
[Depreciated and replaced by V3] This pre-V3 paper is replaced by the corresponding V3 clean-room reconstruction: There Is No Nothing: A Premise-Free Operational Foundation and an Open Verification Platform for Smithian Fold Theory. The V3 source platform is https://github.com/MettaMazza/ernos-labs-sft-platform. The original DOI, concept DOI, version number and files are preserved for transparent historical provenance; this record must not be presented or cited as current V3 work.A comprehensive, highly rigorous consolidated manuscript dismantling black-box AI through the deterministic Smithian Fold Theory. We present exact zero-parameter derivations of the fine-structure constant (137.03599917718), Levinthal's paradox, structural genetics, and SOTA empirical competitive parity in Chess, Symmetric Go, and Natural Language Processing. Unison AI operates at 57 million times the computational efficiency of modern Transformers, tracing physical geometry without gradient descent.
Zico Junius Fernando, Mas Putra Zenno Januarsyah, Firdaus Arifin, Vidyadhara Prawiratama Nugraha · 5 authors
Metaverse has transformed virtual assets into economically valuable objects that challenge conventional concepts of property under Indonesian private law. Although virtual assets such as cryptoassets, non-fungible tokens (NFTs), and metaverse property are widely traded, their legal status remains uncertain, creating ambiguity regarding ownership, transfer, and legal protection. This study examines the normative basis for recognizing virtual assets as objects of property rights within Indonesia's civil law system. Using a normative juridical method with a comparative approach, the study analyzes Indonesian private law alongside developments in England and Wales, Singapore, Japan, and the European Union. The findings demonstrate that virtual assets satisfy the defining characteristics of intangible property, including identifiability, exclusive control, transferability, and economic value, making them capable of recognition as objects of proprietary rights. The study further argues that blockchain-based transfers and smart contracts can operate as legally valid mechanisms for transferring ownership when supported by appropriate legal recognition. To strengthen legal certainty, Indonesia should recognize virtual assets as a distinct category of intangible property, adapt property law to digital transactions, strengthen proprietary remedies, and modernize dispute resolution and cross-border enforcement. These reforms would provide a coherent legal framework for protecting virtual assets and support the development of Indonesia's digital economy.
Abstract: The article proposes "Lagoon Wallet" as a public-learning prototype for blue carbon in transitional waters, using the Venice Lagoon as its case. Through a playable ledger interface, the project translates ecological indicators such as water transparency, disturbance frequency, and seagrass coverage into traceable public entries. Methodologically, the study develops a three-layer mapping structure of indicator, state, and mechanism, and combines it with lightweight user testing (N = 12) to evaluate causal retelling, ledger consultation, and recognition of unsettled boundaries. What it adds is a portable ledger interface, a structured method for translating ecological indicators into public design, and a reusable lightweight evaluation framework. The prototype helps make blue carbon discussable without turning ecological uncertainty into a false sense of settlement.
Fahd Ghalib Basheikh, Ida Widianingsih, Ahmad Zaini Miftah
Decentralized government units in the Global South frequently experience ineffective service delivery because of inadequate funding and weak administrative structures. Using Lamu County Government that allocates bursary funds yet continues to experience operational inefficiencies, this study examines how administrative capacity influences the governance effectiveness of the Lamu County Bursary Programme (LCBP). Guided by Administrative Capacity Theory, the study uses an explanatory sequential mixed methods design using quantitative data from 350 beneficiaries and qualitative data from key informant interviews and focus group discussions. Linear regression results show that administrative capacity is a statistically significant predictor of governance effectiveness (β = 0.627, p < 0.001). Thematic analysis from qualitative data shows three constraints: verification problems, aggravated by geographic dispersion and staffing problems; procedural uncertainty and communication problems, that erode the trust of applicants; and a structural timing penalty, where administrative delays reduce the timeliness and reliability of bursary support, sometimes resulting in temporary school exclusion. The results indicate that the LCBP experiences a capability trap, formal structures are in place but service delivery is weak. Therefore, decentralized units require both financial allocations and effective administrative capabilities. To improve policy outcomes, findings suggest the importance of digitization, staffing at ward level and synchronization of the disbursement calendar with academic cycles.
Data visibility is more vital and decisive than ever before in the current data-driven world of technology. There is a significant upsurge in businesses leveraging digital technology, which has led to a greater amount of data being available than ever before. Additionally, managing the visibility in compliance with the organization's rules and regulations is crucial. The implementation of efficient data visibility will not merely improve decision-making but also streamline business processes with enhanced security. Numerous technologies offer solutions to manage data visibility, and distributed ledger technology (DLT) is one of them. DLT facilitates the execution of different methodologies to strengthen the governance of data visibility in enterprise-grade applications. On the other hand, these DLTs raise concerns regarding data visibility in this decentralized network, as not every enterprise-grade application requires data transparency across all the nodes. In this paper, a detailed systematic review is conducted with a clear focus on two essential data visibility parameters, Access control and anonymity, for the period 2020-2025, following a standardized Preferred Reporting Items for Systematic Review and Meta-Analyses -based breakdown of the selection process. Three clear dimensions of in-depth analysis are presented in the study: first, investigating how DLT can maintain transparency and decentralization in enterprise-grade applications; second, ensuring secure data access management for effective data governance; and third, the approach for anonymization to ensure privacy and security. The key finding highlights the credence of hyperledger fabric, a permissioned DLT, compared to other DLTs and exponentially growing concerns related to data visibility, as well as the conceptual and empirical research contributions made thus far. The limitations presented in this paper formulate a strong basis for research and enhancement of the existing models to offer controlled yet transparent data visibility.
To make the payment system robust and user friendly, decentralized based Scan and Pay system need to be designed. This paper integrates the Unified Payments Interface (UPI) of India with the Solana-based Blockchain to make the payment system decentralized. Solana offers a high throughput and low-cost based decentralized infrastructure which is combined with the simple and reliable UPI system. So, the proposed system enables cryptocurrency transactions linked to UPI while maintaining user friendliness, scalability, and regulatory compliance. The designed method uses a secure architecture powered by smart contracts and modular design. It offers a viable bridge between centralized financial networks and emerging Web3 ecosystems. Proposed Solana-based UPI is compared with the Non-Solana based UPI which is using Blockchain. Results show that there is improvement of 91% in transaction latency and 95% in transaction cost as compared to the Non-Solana based UPI system.
Thandile Nododile, Ayinde M. Usman, Clement N. Nyirenda
Private blockchain networks run with fixed node configurations that cannot adapt to changing workload conditions. Too many nodes serving a light workload waste resources; too few nodes facing heavy demand slow block production and degrade finalisation. The right validator count is hard to determine, as it depends on overlapping factors that shift over time. This paper presents a Takagi-Sugeno (TS) fuzzy inference system that reads live blockchain parameters (block production time, block size, and active node count) and outputs a continuous efficiency score alongside a scaling recommendation: Scale Up, Maintain, or Scale Down. The controller uses triangular membership functions across three linguistic variables, evaluated through a complete 27-rule base with product t-norm aggregation. A key contribution is an empirical recalibration of the membership functions, anchoring linguistic terms to the observed operating range of the testbed rather than to theoretical extremes. The system is evaluated on a 10-node Substrate blockchain network storing real smart water meter data hashes from the Queensland Government open data portal. Statistical analysis across configurations of 4, 7, and 10 active nodes confirms that the controller produces distinct operational profiles reflecting each configuration's provisioning state. In closed-loop experiments, the controller autonomously adjusts validator participation in both directions, activating validators under rising load and removing them under over-provisioning, converging to the same stable equilibrium from both directions. Compared against three threshold-based baselines, it shows fewer scaling oscillations while maintaining comparable block production times. Results show that TS fuzzy inference can support autonomous validator management in private blockchain deployments, with stable scaling behaviour threshold approaches cannot match.
Blockchain systems are undergoing a fundamental transition from decentralized ledgers for digital assets to general-purpose trust infrastructures for verifiable computation, decentralized physical resources, and automated infrastructure management. Meanwhile, the limitations of the Blockchain as a Service (BaaS) model stem from a common structural problem: outsourcing control of infrastructure to third-party service providers inevitably involves a systemic surrender of trust, flexibility, and data sovereignty. RISC-V, with its open, modular, and extensible design, provides a general-purpose computing foundation for public blockchains that is open, low-level, compileable, verifiable, and scalable. Inspired by the development and characteristics of eSIM, the embedded Blockchain infrastructure management (eBIM) is defined as a software-hardware collaborative paradigm for blockchain infrastructure management with RISC-V. This study aims to provide a comprehensive survey on eBIM supporting research and technologies, to answer the following research questions (RQs): RQ1 What is eBIM? RQ2 How does eBIM work? RQ3 What can eBIM do? By introducing the concept of eBIM, this paper establishes a foundational reference for researchers, hardware architects, and protocol designers in this rapidly evolving landscape, including cryptographic acceleration, trusted execution environments, zero-knowledge virtual machines, and smart contract execution engines. The prospects of the proposed e-BIM and its future research directions are indicated in this paper.
Electronic voting must keep individual ballots private while letting anyone verify the final tally. This paper presents an architecture that meets both goals without a trusted key dealer: each voter encrypts a ballot in the browser with a self-generated secret key under the Paillier additive homomorphic cryptosystem, and no party ever holds every key. Two server roles divide the tally. A collector combines the voters' per-ballot auxiliary values into a single group element; an aggregator uses that element to cancel the voters' random masks inside the homomorphic product and recover the exact vote sum, learning the result but no individual ballot. The Solana blockchain records every ciphertext immutably and enforces the election lifecycle, while a native C library (libtommath) performs the heavy modular arithmetic. We state six assumptions under which the protocol is correct and prove product homomorphism, mask cancellation, and sum recovery; privacy rests on the Decisional Composite Residuosity (DCR) assumption for the additive layer together with a Diffie-Hellman-style assumption on the masking base. A bit-packing scheme places an entire multi-candidate ballot in one ciphertext, cutting client work, on-chain transactions, storage, and tally cost by a factor of k (the candidate count); the slot width b is free, with only k*b bounded by log_2(N). With b = 25 and a 255-bit modulus the scheme supports ten candidates and up to 2^25 - 1 = 33,554,431 votes per candidate, about 335 million ballots, and tallies 50,000 ballots in under one second. Finally, running the collector and aggregator inside attested secure enclaves makes the tally tamper-resistant and prevents cross-role collusion to deanonymize voters. The proof-of-concept implementation is open-source; a worked numerical example in the appendix reproduces the full pipeline.
Leopold Müller, Jana Elsner, Thomas Niedermayer, Bernhard Haslhofer · 7 authors
Address clustering is an important technique in blockchain forensics, widely employed by law enforcement to trace illicit crypto asset flows. The multi-input heuristic (MIH), which clusters addresses potentially associated with the same entity, is the most widely used. Yet, despite its broad adoption, the MIH has rarely been evaluated against reliable ground truth data. We implement a reusable evaluation framework covering nine established metrics and apply it to ground truth address-to-entity mappings obtained directly from European crypto asset service providers under legally mandated reporting obligations. When evaluation is restricted to reported addresses, the MIH appears strong at dataset level: we observe no mergers between reported services and recover same-service address pairs with recall 0.71. However, this result is driven by one large service and ignores unlabeled addresses absorbed into full clusters. Metrics that assess the full clusters show substantially lower precision and recall (0.36 and 0.44), meaning that services are often only partially recovered or embedded in larger clusters. Entity-level results further reveal near-complete failures for some services. When MIH-based clusters are used to support criminal suspicion, preliminary seizure of crypto assets to secure later forfeiture/ confiscation, or as evidence in trial proceedings, prosecutors and judges must account for the heuristic's metric-dependent and entity-dependent reliability.
Existing cybercrime classification schemas capture contact metadata and financial transactions but omit the psychological manipulation techniques perpetrators employ. We present a forensic schema (four categories, 35 questions) adding 11 manipulation indicators and cryptocurrency evidence fields to established forensic foundations. Applied to 10,994 victim reports via large language model (LLM)-driven annotation and validated against two human annotators (mean LLM-human $κ= 0.69$, matching inter-annotator $κ= 0.68$), the schema revealed a statistically distinct manipulation profile for each major fraud type (Cramer's $V$ up to $0.790$). A rationale-based evidence audit nonetheless exposed a forensic detail gap: detection of manipulation techniques was reliable, but victim narratives varied widely in the actionable detail supporting each Yes answer, and blockchain-specific identifiers were nearly absent. These findings point to AI-assisted victim intake with schema-informed follow-up questions as the most direct way to close the gap. The tiered annotation strategy also provides a reusable template for LLM-based extraction from other forensic text domains.
Nowadays, mobile forensics is less explored in Digital Forensics case analysis due to the increase in data protection mechanisms implemented by tech companies (i.e., Google for Android and Apple for iOS). For example, the physical acquisition or analysis of specific directories under super-user protection would corrupt the evidence; access to such data is protected, and bypassing this protection requires either privilege escalation or custom ROM installation, leading to the modification of the device state. At the same time, the demand for mobile technologies and their respective communication systems is increasing exponentially, exposing numerous security threats and risks. For that reason, this paper presents a Mobile Live Intelligent Forensics Examination (MoLIFE), a novel Digital Forensics (DF) methodology for data acquisition and analysis of mobile devices. The proposed methodology is based on NIST SP800-101 for the DF process. MoLIFE can be integrated with new and emerging technologies by exploiting their power (e.g., AI, blockchain, quantum computing). MoLIFE can also be used to prevent cyber threats and incidents, as well as DF post-mortem analysis, offering examples of applying the MoLIFE methodology and good practices for the future. To prove the technical feasibility of the methodology, a small case study on Android devices data acquisition via the mDT will be presented. As the methodology is based on new and emerging technologies, it depends on their limitations that would be overcome in a few years.
Smart contract compilers are critical to ensuring the correctness of public blockchains whose defining characteristics are open-source and immutable code. We created SolSmith, a semantics-aware differential fuzz testing tool, to improve the quality of the Solidity compiler -- the most popular compiler for the Ethereum blockchain -- and spent over three years finding compiler defects that produce incorrect code. We call these defects miscompilation bugs. During this time period, we have discovered 25 miscompilation bugs that went unnoticed, some for multiple years. Our first contribution is to make compiler testing more rigorous. SolSmith achieves this goal by generating valid test programs that are likely to stress test code generation and optimization components. This helps SolSmith find bugs missed during routine testing that could potentially have serious implications for smart contracts and their users. Our second contribution is a qualitative and quantitative analysis of miscompilation bugs that we found in the Solidity compiler. We classify miscompilation bugs found by SolSmith based on their nature, root-causes, and impact on end-users. This sheds light on some pitfalls of optimizing compilers.
Monero is a privacy-focused cryptocurrency that deploys the Dandelion++ protocol and incorporates anonymity networks (such as Tor and I2P) to prevent malicious attackers from linking transactions with their source IPs. In this paper, we demonstrate that Monero's integration of the Tor network introduces a fundamental vulnerability: a Monero Tor node's originated transactions are exclusively forwarded to two outgoing Tor hidden service nodes (proxy nodes) prior to clearnet propagation, enabling an adversary to capture originated transactions by occupying the target node's outgoing connections. Based on this observation, we propose \textit{ProxyMark}, a three-stage deanonymization framework for the Monero Tor network, comprising node role identification, originated transaction identification, and node location deanonymization. Through experiments on the live Tor network, Monero mainnet, and testnet, we empirically demonstrate the effectiveness of \textit{ProxyMark} in successfully deanonymizing transactions originating from Monero nodes over Tor.
Artificial intelligence (AI) and blockchain are two of the most transformative technologies of our time, each facing distinct challenges. Blockchain struggles with scalability and efficiency, while AI depends on the integrity of the data it consumes. Yet their proximity in the data value chain enables them to complement one another: AI can optimize blockchain systems through fraud detection, smart contract auditing, or enhanced analytics, while blockchain provides AI with secure, verifiable data crucial for accuracy. The technological convergence of AI and blockchain already reshapes industries such as supply chain management, finance, healthcare, energy, and intellectual property. Emerging solutions—ranging from decentralized data infrastructures to autonomous AI agents—illustrate the growing importance of this technological synergy. Companies implementing AI–blockchain solutions demonstrate enhanced performance, new data monetization opportunities, and even revenue growth. However, convergence raises challenges such as interoperability, reliance on trusted oracles, decentralized data inefficiencies, or regulatory uncertainty. This chapter builds on theories of technological convergence and disruptive innovation to assess the potential of AI–blockchain integration. Drawing on case studies and expert insights, it provides practical frameworks and roadmaps for decision-makers aiming to leverage this convergence as a driver of the next wave of digital transformation.