As large language models (LLMs) are used in sensitive fields, accurately verifying their computational provenance without disclosing their training datasets poses a significant challenge, particularly in regulated sectors such as healthcare, which have strict requirements for dataset use. Traditional approaches either incur substantial computational cost to fully verify the entire training process or leak unauthorized information to the verifier. Therefore, we introduce ZKPROV, a novel cryptographic framework allowing users to verify that the LLM's responses to their prompts are trained on datasets certified by the authorities that own them. Additionally, it ensures that the dataset's content is relevant to the users' queries without revealing sensitive information about the datasets or the model parameters. ZKPROV offers a unique balance between privacy and efficiency by binding training datasets, model parameters, and responses, while also attaching zero-knowledge proofs to the responses generated by the LLM to validate these claims. Our experimental results demonstrate sublinear scaling for generating and verifying these proofs, with end-to-end overhead under 3.3 seconds for models up to 8B parameters, presenting a practical solution for real-world applications. We also provide formal security guarantees, proving that our approach preserves dataset confidentiality while ensuring trustworthy dataset provenance.
The decentralized finance (DeFi) community has grown rapidly in recent years, pushed forward by cryptocurrency enthusiasts interested in the vast untapped potential of new markets. The surge in popularity of cryptocurrency has ushered in a new era of financial crime. Unfortunately, the novelty of the technology makes the task of catching and prosecuting offenders particularly challenging. Thus, it is necessary to implement automated detection tools related to policies to address the growing criminality in the cryptocurrency realm.
Software-defined networking (SDN) enhances network management by centralizing control, but its reliance on a single controller introduces vulnerabilities such as distributed denial-of-service (DDoS) attacks and unauthorized access. Traditional SDN architectures lack authentication mechanisms for verifying network compliment, allowing malicious switches to disrupt operations. To mitigate these risks, blockchain technology provided a decentralized and tamper-proof authentication mechanism by recording the registration status of network devices using smart contracts. By integrating blockchain with the SDN controller (RYU) via Web3, switches are authenticated based on their Datapath Identifier (DPID), enabling secure enforcement of OpenFlow policies. This approach ensures that only registered switches can forward traffic, effectively isolating unauthorized devices and reducing the impact of DDoS attacks. The prototype implementation demonstrates improved security and trust in SDN environments. We combined blockchain’s immutability with SDN’s dynamic programmability.
Efficient and secure sharing of scientific data remains a key challenge in the Open Science framework, especially in terms of data authenticity, provenance and privacy. Traditional digital repositories improve access but often lack decentralized mechanisms that guarantee integrity and traceability. Blockchain technology provides a potential solution through tamper-proof records and distributed consensus, while Zero Knowledge Proofs (ZKP) can enhance privacy protection. This study explores how blockchain and ZKP can be integrated for decentralized scientific data management. A systematic literature review reveals limited application of these combined technologies in Open Science, highlighting a research gap and the need for solutions that support transparent, secure and privacy-preserving data sharing in accordance with FAIR principles.
Florian Spychiger, Sabrina Wollenschläger, Matthias Hafner, Nicolas Oderbolz
Decentralized autonomous organizations (DAOs) have gained popularity over the last few years. Many projects use a DAO for community-based decisions and use a token to enable governance processes and foster participation. The setup of these tokens varies from DAO to DAO. While there are some general tokenomics frameworks, there is no DAO-specific framework including designs of multiple tokens. In this short paper, we aim at exploring the development of such DAO tokenomics framework. To unravel requirements and benefits of such a framework, we conduct interviews with six experts from the Swiss blockchain ecosystem. Switzerland is at the forefront of blockchain development and therefore well suited to serve as an exploration ground. Our results show that a DAO tokenomics framework needs to provide clear guidance while still being flexible to diverse project needs. It may bring along economic gains coupled with a risk reduction and an innovation boost for Switzerland. These benefits could be generalized to other jurisdictions making the development of a DAO tokenomics framework worthwhile.
With the rapid advancement of technology, cloud computing has emerged as the most popular and promising service platform. A cloud user can delegate heavy computation tasks to cloud servers. To ensure the correctness of outsourced processing (e.g., machine learning and data mining), the cloud server must prove that the processing has been executed properly. However, even without malicious intent, it is possible for a cloud server to produce incorrect results. Consequently, clients may outsource the same task to multiple cloud servers and receive various results, aiding them in selecting the best outcome. To protect data privacy, the cloud server must encrypt the results before sending them back to the user. Yet, processing and verifying encrypted results remain significant challenges. To avoid the expensive computational overhead of decrypting ciphertexts from cloud servers one by one, clients prefer to use homomorphic encryption (HE) to obtain the combined output from a single server. However, existing schemes fall short of efficiently verifying the correctness of computations over encrypted data processed by multiple cloud servers, especially in extracting the results computed by each server. In this paper, we introduce a new framework for verifiable outsourced computing systems. In this system, each cloud server's computation result is protected by Paillier encryption, and the edge server can verify these results using zero-knowledge proofs and aggregate the verified ciphertexts. The client can extract the combined plaintext through the Base-3 conversion algorithm to identify each cloud server's results and any non-participating servers. We also prove the security of our scheme and analyze its performance from both theoretical and experimental aspects. Performance analysis shows that our system significantly reduces the client's workload and is userfriendly
The recent EU regulation on Markets in Crypto Assets Regulation (MiCA) represents significant progress in establishing a multi-jurisdiction framework for crypto-assets that will enable the greater participation of consumers in the digital assets industry. One type of token recognized by MiCA is the asset-referenced token, where the value-bearing physical asset being referenced by the token is external to the token. We discuss several design considerations for the on-chain and off- chain metadata for the asset-referenced token that represents physical real-world assets. The EU Data Spaces provides an interesting data management framework for the off-chain metadata underpinning the MiCA asset-referenced tokens, including asset definition schemas, asset profiles, digitized asset records, and tokenized asset records. The Web3 decentralized registries for assets-related metadata should be a promising application of the data spaces framework in the EU.
Independent Researcher, USA, Damodar Bihani, Bright Chibunna Ubamadu, Signal Alliance Technology Holding, Nigeria · 6 authors
The integration of blockchain technology into the tokenization of real-world assets (RWAs) is revolutionizing how value is stored, transferred, and accessed globally. This paper proposes a scalable framework for cross-functional collaboration in Web3 product development focused on blockchain-based tokenized RWAs. Tokenization enables physical assets such as real estate, commodities, and intellectual property to be digitized into blockchain-based tokens, allowing for fractional ownership, increased liquidity, and enhanced accessibility. However, the successful development and deployment of such Web3 products require an interdisciplinary approach that combines technological innovation, legal compliance, financial modeling, and user experience design. Our framework addresses these needs by enabling seamless collaboration between developers, legal experts, financial analysts, and UX/UI designers throughout the product lifecycle. We present a modular architecture built on interoperable blockchain protocols such as Ethereum and Polkadot, integrating smart contracts, decentralized identifiers (DIDs), and oracles for real-time asset verification. The framework emphasizes agile product development practices and leverages decentralized autonomous organization (DAO) structures to facilitate decision-making and community governance. Furthermore, we explore how regulatory-compliant token standards, such as ERC-1400, can be incorporated to ensure adherence to jurisdiction-specific asset ownership and transfer laws. This study includes a case analysis of cross-functional product teams building tokenized real estate platforms and carbon credit marketplaces, demonstrating how scalable collaboration can accelerate time-to-market and improve transparency, trust, and user adoption. Our findings highlight that such a collaborative framework significantly reduces technical debt and improves legal and financial risk mitigation. The framework also enhances stakeholder alignment through integrated project management tools and on-chain documentation. By offering a structured, scalable, and adaptable approach, this framework positions Web3 product teams to unlock the full potential of tokenized RWAs in a decentralized economy. It serves as a critical guide for developers, entrepreneurs, regulators, and investors aiming to leverage blockchain technology in building trustworthy, scalable, and cross-functional Web3 applications.
Sajan Poudel, Rasib Khan, Aalok Dhonju, Nishar Miya
Edge computing is revolutionizing digital infrastructures by enabling low-latency processing, bandwidth optimization, and real-time decision-making across applications like IoT, smart cities, and industrial automation. However, its decentralized nature introduces significant security challenges, making trust and accountability crucial for reliable service delivery. Secure provenance management is vital to address these challenges, ensuring that distributed services and high data volumes are protected from tampering and unauthorized access. This research presents SPHERE, a scalable framework that integrates distributed ledgers, digital signatures, and cryptographic techniques to ensure service integrity and traceability. By leveraging blockchain for tamper-proof provenance, EdgeX Foundry for service orchestration, and an off-chain database for load balancing, our approach enhances security while maintaining system performance. Empirical analysis through a proof-of-concept deployment on a virtualized testbed validates its effectiveness in strengthening service reliability, auditability, and compliance, addressing critical gaps in edge service security and underscoring the importance of secure provenance-awareness in decentralized environments.
Alexandra Vultureanu‐Albişi, Costin Bădică, Mirjana Ivanović
The Internet of Things (IoT) paradigm is evolving and the Next-Generation IoT (NG-IoT) ecosystem will incorporate distributed ledger and blockchain technology, AI-adapted components, and intelligent edge solutions that take advantage of edge computing, Artificial Intelligence (AI), networks, and communications. In addition to the low integration of eXplainable Artificial Intelligence (XAI) in the IoT or NG-IoT contexts, the explainability of these systems is rarely evaluated. Due to these limitations, we thoroughly examined the current state of XAI integration with IoT services. We propose a new conceptual framework called eXING-IoT (eXplainability Integrated in the Next Generation IoT) for better NG-IoT systems' explainability integration and evaluation. This includes a list of qualities that future NG-IoT environments should have, thus paving the way for the advancement of NG-IoT beyond the state of the art.
Open access
Explainable Artificial Intelligence (XAI)
Scientific Computing and Data Management
Artificial Intelligence in Healthcare and Education
Critical peer review of scientific manuscripts presents a significant challenge for Large Language Models (LLMs), partly due to data limitations and the complexity of expert reasoning. This report introduces Persistent Workflow Prompting (PWP), a potentially broadly applicable prompt engineering methodology designed to bridge this gap using standard LLM chat interfaces (zero-code, no APIs). We present a proof-of-concept PWP prompt for the critical analysis of experimental chemistry manuscripts, featuring a hierarchical, modular architecture (structured via Markdown) that defines detailed analysis workflows. We develop this PWP prompt through iterative application of meta-prompting techniques and meta-reasoning aimed at systematically codifying expert review workflows, including tacit knowledge. Submitted once at the start of a session, this PWP prompt equips the LLM with persistent workflows triggered by subsequent queries, guiding modern reasoning LLMs through systematic, multimodal evaluations. Demonstrations show the PWP-guided LLM identifying major methodological flaws in a test case while mitigating LLM input bias and performing complex tasks, including distinguishing claims from evidence, integrating text/photo/figure analysis to infer parameters, executing quantitative feasibility checks, comparing estimates against claims, and assessing a priori plausibility. To ensure transparency and facilitate replication, we provide full prompts, detailed demonstration analyses, and logs of interactive chats as supplementary resources. Beyond the specific application, this work offers insights into the meta-development process itself, highlighting the potential of PWP, informed by detailed workflow formalization, to enable sophisticated analysis using readily available LLMs for complex scientific tasks.
Ethereum is a decentralized blockchain system that relies on a peer-to-peer (P2P) network. Understanding the topology of this P2P network is crucial for assessing the security, reliability, and user anonymity of Ethereum (ETH). However, the routing table of Ethereum network nodes is private and inaccessible, making the Ethereum network topology hidden. Safeguarding the topology is essential to protecting Ethereum's privacy and security. Therefore, attempting to measure the entire ETH network's topology is a crucial foundation. This paper proposes a completely passive method called ETHNetPRecover, which only requires monitoring transactions to recover the entire ETH network's topology. This method can determine the presence of an edge between two nodes with 96.8% precision. We compared our method's performance with similar approaches and found it recovers the Ethereum network topology faster without additional transaction costs. Using the recovered topology, we tracked and marked transaction entry nodes in the Ethereum Mainnet over seven days. We discovered that 93% of the transactions during this period entered the Ethereum Mainnet network through 30 entry nodes. Through our evaluation, we found that our method has high precision, is efficient, and incurs no additional costs, making it a relatively effective approach.
Many educational institutions worldwide now use blockchain to verify electronic document, often relying on Ethereum 1.0, which uses proof of work (PoW) or proof of authority (PoA). However, Ethereum 2.0, launched in 2022 by Ethereum Foundation operates on proof of stake (PoS). This study provides comparative analysis of PoS and PoA consensus in Ethereum environment specifically focusing on performance and scalability in the context of academic transcript databases. To demonstrate this, a student academic reputation information system was developed using two different blockchain technologies: Ethereum 1.0 with PoA and Ethereum 2.0 with PoS. This setup was used to obtain comparative analysis data for the two blockchain systems by measuring the throughput and latency. We observed how these platforms responded to an increasing number and frequency of transactions with Hyperledger Caliper. Results indicates that in performance testing, both consensus mechanisms exhibited. Scalability tests revealed that both consensus mechanisms experienced increased latency with higher loads. However, PoA system was superior in average throughput and latency than PoS system except in high transaction of data addition. The experiment result show that PoA system better than PoS system in context of academic transcript databases, making it more suitable to be implemented on that context.
The unprecedented growth of digital health ecosystems, fueled by electronic health records (EHRs), wearable devices, telemedicine, and AI-driven diagnostics, has amplified the critical need for reliable data provenance mechanisms. Provenance, defined as the comprehensive history of data generation, access, transformation, and transfer, ensures that stakeholders—including patients, clinicians, insurers, researchers, and regulators—can trust the authenticity, integrity, and accountability of healthcare information. Traditional provenance systems, often centralized, are vulnerable to insider manipulation, cyberattacks, data silos, and audit inefficiencies, thereby undermining trust and regulatory compliance. Distributed Ledger Systems (DLS), encompassing blockchain, permissioned ledgers, and Directed Acyclic Graphs (DAGs), offer a paradigm shift by enabling immutable, transparent, and tamper-evident provenance trails across diverse healthcare stakeholders. This manuscript provides an in-depth exploration of DLS-enabled healthcare data provenance by reviewing current literature, identifying research gaps, and developing a methodological framework tested through simulated experiments. Empirical evaluation demonstrates that distributed ledgers reduce provenance validation time by 57–71%, accelerate audit processes by up to 70%, and significantly enhance regulatory traceability under HIPAA and GDPR requirements. Moreover, patient-centric smart contracts and decentralized identifiers foster individual ownership and interoperability, reshaping data governance models toward inclusivity and transparency. While challenges such as scalability, energy efficiency, and privacy-preserving erasure remain, the findings highlight DLS as a transformative infrastructure for establishing trustworthy healthcare ecosystems. The study concludes by recommending hybrid ledger architectures, cryptographic privacy enhancements, and supportive policy frameworks to ensure sustainable, ethical, and globally interoperable healthcare data provenance systems.
Cheick Tidiane Bâ, Benjamin A. Steer, Matteo Zignani, Richard G. Clegg
Blockchain technology and cryptocurrencies have garnered considerable attention over the past 15 years. The term Web3 (sometimes Web 3.0) has been coined to define a possible direction for the web based on the use of decentralisation via blockchain. Cryptocurrencies are characterised by high market volatility and susceptibility to substantial crashes, issues that require temporal analysis methodologies able to tackle the high temporal resolution, heterogeneity, and scale of blockchain data. While existing research attempts to analyse crash events, fundamental questions persist regarding the optimal timescale for analysis, differentiation between long-term and short-term trends, and the identification and characterisation of shock events within these decentralised systems. This article addresses these issues by examining cryptocurrencies traded on the Ethereum blockchain, with a spotlight on the crash of the stablecoin TerraUSD (UST) and the currency LUNA designed to stabilise it. Utilising complex network analysis and a multi-layer temporal graph allows the study of the correlations between the layers representing the currencies and system evolution across diverse timescales. The investigation sheds light on the strong interconnections among stablecoins pre-crash and the significant post-crash transformations. We identify anomalous signals before, during, and after the collapse, emphasising their impact on graph structure metrics and user movement across layers. This article is novel in its use of temporal, cross-chain graph analysis to explore a cryptocurrency collapse. It emphasises the importance of temporal analysis for studies on web-derived data. In addition, the methodology shows how graph-based analysis can enhance traditional econometric results. Overall, this research carries implications beyond its field, for example, for regulatory agencies aiming to safeguard users could use multi-layer temporal graphs as part of their suite of analysis tools.
Alex Veith, Patrick R. Carney, Aiqing Wu, Brenda L. Rojas · 9 authors
Scientific progress benefits from the sharing of "research assets" such as data, reagents, models, and experimental samples. To improve asset shareability, we evaluated the availability, quality, and characterization of recombinant DNA molecules, recombinant mouse models, and tissue samples described by our laboratory in ten publications spanning over thirty years. Employing state-of-the-art molecular technologies, we identified existing samples, updated their localization, generated modern sequences and maps of recombinant models, and ported the associated metadata to an internal blockchain-dependent resource using a standardized description for each asset class. We also created non-fungible tokens representing research assets on the public blockchain network Solana. In addition to providing an audit of previously reported shareable assets and improving the value of recombinant models, this re-analysis also provides evidence for the utility of extant tissue samples that may be difficult and expensive to regenerate. The results demonstrate how retrospective analysis can improve and expand upon the spectrum of shareable research assets through updates on molecular characterization and physical location, as well as improving the availability of biological samples of potential high experimental value. Moreover, the development of a decentralized ledger harboring this revised metadata provides a path to the description and tokenization of scientific assets and provides a strategy to extend the life of scientific assets even after laboratories or sources close.
Jens Ernstberger, Jan Lauinger, Yulin Wu, Arthur Gervais · 5 authors
Transport Layer Security (TLS) is foundational for safeguarding client-server communication. However, it does not extend integrity guarantees to third-party verification of data authenticity. If a client wants to present data obtained from a server, it cannot convince any other party that the data has not been tampered with. TLS oracles ensure data authenticity beyond the client-server TLS connection, such that clients can obtain data from a server and ensure provenance to any third party, without server-side modifications. Generally, a TLS oracle involves a third party, the verifier, in a TLS session to verify that the data obtained by the client is accurate. Existing protocols for TLS oracles are communication-heavy, as they rely on interactive protocols. We present ORIGO, a TLS oracle with constant communication. Similar to prior work, ORIGO introduces a third party in a TLS session, and provides a protocol to ensure the authenticity of data transmitted in a TLS session, without forfeiting its confidentiality. Compared to prior work, we rely on intricate details specific to TLS 1.3, which allow us to prove correct key derivation, authentication and encryption within a Zero Knowledge Proof (ZKP). This, combined with optimizations for TLS 1.3, leads to an efficient protocol with constant communication in the online phase. Our work reduces online communication by 375× and online runtime by up to 4.6×, compared to prior work.
Smart contracts, predominantly written in Solidity and executed on blockchains like Ethereum, are immutable, making functional correctness paramount: once deployed, bugs and vulnerabilities become permanent. Despite rapid progress in transformer-based code LLMs, existing evaluations of Solidity code completion rely heavily on surface-form metrics (e.g., BLEU, CrystalBLEU) or hand-grading, which poorly correlate with functional correctness. Unlike Python, Solidity lacks large-scale and execution-based benchmarks, hindering systematic assessment and optimization of LLMs for smart contract development. To bridge this research gap, we introduce SolBench, a comprehensive benchmark and automated testing pipeline for Solidity, designed to emphasize functional correctness via differential fuzzing. SolBench contains 28,825 functions from 7,604 contracts collected from Etherscan (genesis to 2024), spanning 10 popular domains. We benchmark 14 diverse LLMs (open/closed, 1.3B to 671B parameters, general/code-specific, with/without reasoning). The dominant failure mode is missing crucial details (e.g., type definitions, state variables) in intra-contract context. Providing full-contract context mitigates this and improves code completion accuracy. However, full-context inference can be prohibitively expensive in practice. Generating outputs with large context windows using state-of-the-art models often incurs significant costs, rendering naive context scaling economically impractical. Crucially, most of a contract is irrelevant to implementing a given function; only a small subset of details is needed. To exploit this, we propose Retrieval-Augmented Repair (RAR), which integrates retrieval into code repair: it uses the executor's error messages to extract only the most relevant snippets from the full contract. RAR sharply reduces input length for function completion, improving accuracy while significantly cutting computational cost. We further analyze retrieval and code repair strategies within RAR, showing substantial improvements in accuracy and efficiency. SolBench and our RAR framework enable principled evaluation and cost-effective improvement of Solidity code generation. Dataset and code are available at https://github.com/ZaoyuChen/SolBench.
Chhavi Yadav, Evan Monroe Laufer, Dan Boneh, Kamalika Chaudhuri
In principle, explanations are intended as a way to increase trust in machine learning models and are often obligated by regulations. However, many circumstances where these are demanded are adversarial in nature, meaning the involved parties have misaligned interests and are incentivized to manipulate explanations for their purpose. As a result, explainability methods fail to be operational in such settings despite the demand \cite{bordt2022post}. In this paper, we take a step towards operationalizing explanations in adversarial scenarios with Zero-Knowledge Proofs (ZKPs), a cryptographic primitive. Specifically we explore ZKP-amenable versions of the popular explainability algorithm LIME and evaluate their performance on Neural Networks and Random Forests. Our code is publicly available at https://github.com/emlaufer/ExpProof.
This paper introduces 3MEthTaskforce (https://3meth.github.io), a multi-source, multi-level, and multi-token Ethereum dataset addressing the limitations of single-source datasets. Integrating over 300 million transaction records, 3,880 token profiles, global market indicators, and Reddit sentiment data from 2014-2024, it enables comprehensive studies on user behavior, market sentiment, and token performance. 3MEthTaskforce defines benchmarks for user behavior prediction and token price prediction tasks, using 6 dynamic graph networks and 19 time-series models to evaluate performance. Its multimodal design supports risk analysis and market fluctuation modeling, providing a valuable resource for advancing blockchain analytics and decentralized finance research.
Rashid Ul Haq, Rahim Khan, Fahad Alturise, Shafrida Sahrani · 6 authors
Recent technological advances have enabled researchers to investigate various novel approaches utilized to manage allograft transplants and overcome the challenges of conventional centralized systems. The rising need for transparency, efficiency, and, especially, security in this highly sensitive medical procedure necessitates the use of decentralized solutions like blockchain rather than existing centralized approaches. However, the current state of research is theoretical and unproven, and allograft management lacks any reliable, cost-effective, or data-proven solution. In this paper, we propose an Ethereum blockchain-based allograft transplantation management system that can address all of those issues linked to the existing solutions. The proposed approach aims to enhance traceability, transparency, and data provenance across the entire allograft transplant process. We present six reliable and cost-efficient algorithms, as well as a comprehensive system architecture, to provide valuable insight into system implementation complexity. We have designed an efficient smart contract implementing the proposed algorithms to ensure flawless execution of allograft donation, transportation, and transplantation. We conduct thorough tests, validation, security, cost, throughput, and latency assessments of the system in order to contrast its effectiveness with existing solutions and results shows that our solution is cost-effective, as well as secure and efficient. We generalized the proposed solution so that, with minimal changes, it could be used for other problems and addressed some of the technical and ethical challenges.
The reproducibility of scientific simulations is one of the key challenges of scientific research. Current best practices involve version-controlled code, tracking dependencies, specifying hardware configurations, and sometimes using Docker containers to enable one-click simulation setups. However, these approaches still fall short of achieving true reproducibility. For example, Docker depends on the underlying host kernel, and high-performance computing (HPC) codes often link with specific kernel modules and headers. Over time, changes in host kernel versions can render Dockerized simulations unusable. Furthermore, non-deterministic simulations, such as Monte Carlo methods, may not yield identical results even when rerun on the same hardware with the same code.This talk explores the potential of blockchain technology to address these challenges. By running simulations natively on-chain (via smart contracts) and emitting logs of each state transition, we can achieve reproducibility while also verifying the simulation's authenticity (associating the original author of the simulation and the reporting author).Other potential ideas include using zero-knowledge proofs to hash the call stack and the stack memory into a Merkle tree or also to think about the tokenisation of compute.We will delve into the technical feasibility and potential benefits of this approach, including its implications for trust, transparency, and the future of scientific research.