Blockchain Papers

Follow blockchain research across journals, conferences, and preprint repositories.

300 papersLast indexed Aug 31, 2026
Search papers

Paper index

300 results · page 6 of 13

Clear filters
Jun 26, 2025·arXiv (Cornell University)
0 cites
ZKPROV: A Zero-Knowledge Approach to Dataset Provenance for Large Language Models

Mina Namazi, Alexander Nemecek, Erman Ayday

As large language models (LLMs) are used in sensitive fields, accurately verifying their computational provenance without disclosing their training datasets poses a significant challenge, particularly in regulated sectors such as healthcare, which have strict requirements for dataset use. Traditional approaches either incur substantial computational cost to fully verify the entire training process or leak unauthorized information to the verifier. Therefore, we introduce ZKPROV, a novel cryptographic framework allowing users to verify that the LLM's responses to their prompts are trained on datasets certified by the authorities that own them. Additionally, it ensures that the dataset's content is relevant to the users' queries without revealing sensitive information about the datasets or the model parameters. ZKPROV offers a unique balance between privacy and efficiency by binding training datasets, model parameters, and responses, while also attaching zero-knowledge proofs to the responses generated by the LLM to validate these claims. Our experimental results demonstrate sublinear scaling for generating and verifying these proofs, with end-to-end overhead under 3.3 seconds for models up to 8B parameters, presenting a practical solution for real-world applications. We also provide formal security guarantees, proving that our approach preserves dataset confidentiality while ensuring trustworthy dataset provenance.

Open access
2 source records
cs.CR
cs.AI
cs.LG
Original source
Jun 17, 2025·arXiv (Cornell University)
0 cites
Explain First, Trust Later: LLM-Augmented Explanations for Graph-Based Crypto Anomaly Detection

Watson, Adriana, Richards, Grant, Schiff, Daniel

The decentralized finance (DeFi) community has grown rapidly in recent years, pushed forward by cryptocurrency enthusiasts interested in the vast untapped potential of new markets. The surge in popularity of cryptocurrency has ushered in a new era of financial crime. Unfortunately, the novelty of the technology makes the task of catching and prosecuting offenders particularly challenging. Thus, it is necessary to implement automated detection tools related to policies to address the growing criminality in the cryptocurrency realm.

Open access
2 source records
cs.CE
cs.AI
cs.CR
Original source
Jun 6, 2025·DOAJ (DOAJ: Directory of Open Access Journals)
0 cites
Trust, Privacy and Authenticity in Scientific Data Sharing

Almeida, Joana, Santos, Rita, Martins, Ciro, Gomes, Hélder · 9 authors

Efficient and secure sharing of scientific data remains a key challenge in the Open Science framework, especially in terms of data authenticity, provenance and privacy. Traditional digital repositories improve access but often lack decentralized mechanisms that guarantee integrity and traceability. Blockchain technology provides a potential solution through tamper-proof records and distributed consensus, while Zero Knowledge Proofs (ZKP) can enhance privacy protection. This study explores how blockchain and ZKP can be integrated for decentralized scientific data management. A systematic literature review reveals limited application of these combined technologies in Open Science, highlighting a research gap and the need for solutions that support transparent, secure and privacy-preserving data sharing in accordance with FAIR principles.

Open access
Research Data Management Practices
Scientific Computing and Data Management
Blockchain Technology Applications and Security
Original source
Jun 5, 2025·2025 Crypto Valley Conference (CVC)
0 cites
Short Paper: Requirements and Benefits of a DAO Tokenomics Framework

Florian Spychiger, Sabrina Wollenschläger, Matthias Hafner, Nicolas Oderbolz

Decentralized autonomous organizations (DAOs) have gained popularity over the last few years. Many projects use a DAO for community-based decisions and use a token to enable governance processes and foster participation. The setup of these tokens varies from DAO to DAO. While there are some general tokenomics frameworks, there is no DAO-specific framework including designs of multiple tokens. In this short paper, we aim at exploring the development of such DAO tokenomics framework. To unravel requirements and benefits of such a framework, we conduct interviews with six experts from the Swiss blockchain ecosystem. Switzerland is at the forefront of blockchain development and therefore well suited to serve as an exploration ground. Our results show that a DAO tokenomics framework needs to provide clear guidance while still being flexible to diverse project needs. It may bring along economic gains coupled with a risk reduction and an innovation boost for Switzerland. These benefits could be generalized to other jurisdictions making the development of a DAO tokenomics framework worthwhile.

Open access
Data Stream Mining Techniques
Scientific Computing and Data Management
Blockchain Technology Applications and Security
Original source
May 29, 2025·Engineering and Technology Journal
0 cites
Blockchain for Tokenized Real-World Assets: A Scalable Framework for Cross-Functional Collaboration in Web3 Product Development

Independent Researcher, USA, Damodar Bihani, Bright Chibunna Ubamadu, Signal Alliance Technology Holding, Nigeria · 6 authors

The integration of blockchain technology into the tokenization of real-world assets (RWAs) is revolutionizing how value is stored, transferred, and accessed globally. This paper proposes a scalable framework for cross-functional collaboration in Web3 product development focused on blockchain-based tokenized RWAs. Tokenization enables physical assets such as real estate, commodities, and intellectual property to be digitized into blockchain-based tokens, allowing for fractional ownership, increased liquidity, and enhanced accessibility. However, the successful development and deployment of such Web3 products require an interdisciplinary approach that combines technological innovation, legal compliance, financial modeling, and user experience design. Our framework addresses these needs by enabling seamless collaboration between developers, legal experts, financial analysts, and UX/UI designers throughout the product lifecycle. We present a modular architecture built on interoperable blockchain protocols such as Ethereum and Polkadot, integrating smart contracts, decentralized identifiers (DIDs), and oracles for real-time asset verification. The framework emphasizes agile product development practices and leverages decentralized autonomous organization (DAO) structures to facilitate decision-making and community governance. Furthermore, we explore how regulatory-compliant token standards, such as ERC-1400, can be incorporated to ensure adherence to jurisdiction-specific asset ownership and transfer laws. This study includes a case analysis of cross-functional product teams building tokenized real estate platforms and carbon credit marketplaces, demonstrating how scalable collaboration can accelerate time-to-market and improve transparency, trust, and user adoption. Our findings highlight that such a collaborative framework significantly reduces technical debt and improves legal and financial risk mitigation. The framework also enhances stakeholder alignment through integrated project management tools and on-chain documentation. By offering a structured, scalable, and adaptable approach, this framework positions Web3 product teams to unlock the full potential of tokenized RWAs in a decentralized economy. It serves as a critical guide for developers, entrepreneurs, regulators, and investors aiming to leverage blockchain technology in building trustworthy, scalable, and cross-functional Web3 applications.

Open access
Scientific Computing and Data Management
Cloud Computing and Resource Management
Original source
May 24, 2025·Connection Science
1 cites
eXING-IoT conceptual framework for explainability integration in next generation-IoT

Alexandra Vultureanu‐Albişi, Costin Bădică, Mirjana Ivanović

The Internet of Things (IoT) paradigm is evolving and the Next-Generation IoT (NG-IoT) ecosystem will incorporate distributed ledger and blockchain technology, AI-adapted components, and intelligent edge solutions that take advantage of edge computing, Artificial Intelligence (AI), networks, and communications. In addition to the low integration of eXplainable Artificial Intelligence (XAI) in the IoT or NG-IoT contexts, the explainability of these systems is rarely evaluated. Due to these limitations, we thoroughly examined the current state of XAI integration with IoT services. We propose a new conceptual framework called eXING-IoT (eXplainability Integrated in the Next Generation IoT) for better NG-IoT systems' explainability integration and evaluation. This includes a list of qualities that future NG-IoT environments should have, thus paving the way for the advancement of NG-IoT beyond the state of the art.

Open access
Explainable Artificial Intelligence (XAI)
Scientific Computing and Data Management
Artificial Intelligence in Healthcare and Education
Original source
May 6, 2025·arXiv (Cornell University)
0 cites
AI-Driven Scholarly Peer Review via Persistent Workflow Prompting, Meta-Prompting, and Meta-Reasoning

Evgeny Markhasin

Critical peer review of scientific manuscripts presents a significant challenge for Large Language Models (LLMs), partly due to data limitations and the complexity of expert reasoning. This report introduces Persistent Workflow Prompting (PWP), a potentially broadly applicable prompt engineering methodology designed to bridge this gap using standard LLM chat interfaces (zero-code, no APIs). We present a proof-of-concept PWP prompt for the critical analysis of experimental chemistry manuscripts, featuring a hierarchical, modular architecture (structured via Markdown) that defines detailed analysis workflows. We develop this PWP prompt through iterative application of meta-prompting techniques and meta-reasoning aimed at systematically codifying expert review workflows, including tacit knowledge. Submitted once at the start of a session, this PWP prompt equips the LLM with persistent workflows triggered by subsequent queries, guiding modern reasoning LLMs through systematic, multimodal evaluations. Demonstrations show the PWP-guided LLM identifying major methodological flaws in a test case while mitigating LLM input bias and performing complex tasks, including distinguishing claims from evidence, integrating text/photo/figure analysis to infer parameters, executing quantitative feasibility checks, comparing estimates against claims, and assessing a priori plausibility. To ensure transparency and facilitate replication, we provide full prompts, detailed demonstration analyses, and logs of interactive chats as supplementary resources. Beyond the specific application, this work offers insights into the meta-development process itself, highlighting the potential of PWP, informed by detailed workflow formalization, to enable sophisticated analysis using readily available LLMs for complex scientific tasks.

Open access
Scientific Computing and Data Management
Original source
May 2, 2025·Bulletin of Electrical Engineering and Informatics
1 cites
Comparative analysis of PoS and PoA consensus in Ethereum environment for blockchain based academic transcript systems

Palguno Wicaksono, Puspanda Hatta, Yusfia Hafid Aristyagama

Many educational institutions worldwide now use blockchain to verify electronic document, often relying on Ethereum 1.0, which uses proof of work (PoW) or proof of authority (PoA). However, Ethereum 2.0, launched in 2022 by Ethereum Foundation operates on proof of stake (PoS). This study provides comparative analysis of PoS and PoA consensus in Ethereum environment specifically focusing on performance and scalability in the context of academic transcript databases. To demonstrate this, a student academic reputation information system was developed using two different blockchain technologies: Ethereum 1.0 with PoA and Ethereum 2.0 with PoS. This setup was used to obtain comparative analysis data for the two blockchain systems by measuring the throughput and latency. We observed how these platforms responded to an increasing number and frequency of transactions with Hyperledger Caliper. Results indicates that in performance testing, both consensus mechanisms exhibited. Scalability tests revealed that both consensus mechanisms experienced increased latency with higher loads. However, PoA system was superior in average throughput and latency than PoS system except in high transaction of data addition. The experiment result show that PoA system better than PoS system in context of academic transcript databases, making it more suitable to be implemented on that context.

Open access
Scientific Computing and Data Management
Original source
Apr 3, 2025·Scientific Journal of Artificial Intelligence and Blockchain Technologies
0 cites
Healthcare Data Provenance Using Distributed Ledger Systems

Niharika Singh

The unprecedented growth of digital health ecosystems, fueled by electronic health records (EHRs), wearable devices, telemedicine, and AI-driven diagnostics, has amplified the critical need for reliable data provenance mechanisms. Provenance, defined as the comprehensive history of data generation, access, transformation, and transfer, ensures that stakeholders—including patients, clinicians, insurers, researchers, and regulators—can trust the authenticity, integrity, and accountability of healthcare information. Traditional provenance systems, often centralized, are vulnerable to insider manipulation, cyberattacks, data silos, and audit inefficiencies, thereby undermining trust and regulatory compliance. Distributed Ledger Systems (DLS), encompassing blockchain, permissioned ledgers, and Directed Acyclic Graphs (DAGs), offer a paradigm shift by enabling immutable, transparent, and tamper-evident provenance trails across diverse healthcare stakeholders. This manuscript provides an in-depth exploration of DLS-enabled healthcare data provenance by reviewing current literature, identifying research gaps, and developing a methodological framework tested through simulated experiments. Empirical evaluation demonstrates that distributed ledgers reduce provenance validation time by 57–71%, accelerate audit processes by up to 70%, and significantly enhance regulatory traceability under HIPAA and GDPR requirements. Moreover, patient-centric smart contracts and decentralized identifiers foster individual ownership and interoperability, reshaping data governance models toward inclusivity and transparency. While challenges such as scalability, energy efficiency, and privacy-preserving erasure remain, the findings highlight DLS as a transformative infrastructure for establishing trustworthy healthcare ecosystems. The study concludes by recommending hybrid ledger architectures, cryptographic privacy enhancements, and supportive policy frameworks to ensure sustainable, ethical, and globally interoperable healthcare data provenance systems.

Open access
Scientific Computing and Data Management
Blockchain Technology Applications and Security
Data Quality and Management
Original source
Apr 3, 2025·ACM Transactions on the Web
1 cites
Investigating the Luna-Terra Collapse through the Temporal Multilayer Graph Structure of the Ethereum Stablecoin Ecosystem

Cheick Tidiane Bâ, Benjamin A. Steer, Matteo Zignani, Richard G. Clegg

Blockchain technology and cryptocurrencies have garnered considerable attention over the past 15 years. The term Web3 (sometimes Web 3.0) has been coined to define a possible direction for the web based on the use of decentralisation via blockchain. Cryptocurrencies are characterised by high market volatility and susceptibility to substantial crashes, issues that require temporal analysis methodologies able to tackle the high temporal resolution, heterogeneity, and scale of blockchain data. While existing research attempts to analyse crash events, fundamental questions persist regarding the optimal timescale for analysis, differentiation between long-term and short-term trends, and the identification and characterisation of shock events within these decentralised systems. This article addresses these issues by examining cryptocurrencies traded on the Ethereum blockchain, with a spotlight on the crash of the stablecoin TerraUSD (UST) and the currency LUNA designed to stabilise it. Utilising complex network analysis and a multi-layer temporal graph allows the study of the correlations between the layers representing the currencies and system evolution across diverse timescales. The investigation sheds light on the strong interconnections among stablecoins pre-crash and the significant post-crash transformations. We identify anomalous signals before, during, and after the collapse, emphasising their impact on graph structure metrics and user movement across layers. This article is novel in its use of temporal, cross-chain graph analysis to explore a cryptocurrency collapse. It emphasises the importance of temporal analysis for studies on web-derived data. In addition, the methodology shows how graph-based analysis can enhance traditional econometric results. Overall, this research carries implications beyond its field, for example, for regulatory agencies aiming to safeguard users could use multi-layer temporal graphs as part of their suite of analysis tools.

Open access
Geology and Paleoclimatology Research
Scientific Computing and Data Management
Space Science and Extraterrestrial Life
Original source
Mar 19, 2025·Biochemical Pharmacology
0 cites
Retrospective analysis and decentralized distribution to improve the lifecycle of Ah receptor research assets

Alex Veith, Patrick R. Carney, Aiqing Wu, Brenda L. Rojas · 9 authors

Scientific progress benefits from the sharing of "research assets" such as data, reagents, models, and experimental samples. To improve asset shareability, we evaluated the availability, quality, and characterization of recombinant DNA molecules, recombinant mouse models, and tissue samples described by our laboratory in ten publications spanning over thirty years. Employing state-of-the-art molecular technologies, we identified existing samples, updated their localization, generated modern sequences and maps of recombinant models, and ported the associated metadata to an internal blockchain-dependent resource using a standardized description for each asset class. We also created non-fungible tokens representing research assets on the public blockchain network Solana. In addition to providing an audit of previously reported shareable assets and improving the value of recombinant models, this re-analysis also provides evidence for the utility of extant tissue samples that may be difficult and expensive to regenerate. The results demonstrate how retrospective analysis can improve and expand upon the spectrum of shareable research assets through updates on molecular characterization and physical location, as well as improving the availability of biological samples of potential high experimental value. Moreover, the development of a decentralized ledger harboring this revised metadata provides a path to the description and tokenization of scientific assets and provides a strategy to extend the life of scientific assets even after laboratories or sources close.

Open access
Scientific Computing and Data Management
CCD and CMOS Imaging Sensors
Original source
Mar 7, 2025·Proceedings on Privacy Enhancing Technologies
1 cites
ORIGO: Proving Provenance of Sensitive Data with Constant Communication

Jens Ernstberger, Jan Lauinger, Yulin Wu, Arthur Gervais · 5 authors

Transport Layer Security (TLS) is foundational for safeguarding client-server communication. However, it does not extend integrity guarantees to third-party verification of data authenticity. If a client wants to present data obtained from a server, it cannot convince any other party that the data has not been tampered with. TLS oracles ensure data authenticity beyond the client-server TLS connection, such that clients can obtain data from a server and ensure provenance to any third party, without server-side modifications. Generally, a TLS oracle involves a third party, the verifier, in a TLS session to verify that the data obtained by the client is accurate. Existing protocols for TLS oracles are communication-heavy, as they rely on interactive protocols. We present ORIGO, a TLS oracle with constant communication. Similar to prior work, ORIGO introduces a third party in a TLS session, and provides a protocol to ensure the authenticity of data transmitted in a TLS session, without forfeiting its confidentiality. Compared to prior work, we rely on intricate details specific to TLS 1.3, which allow us to prove correct key derivation, authentication and encryption within a Zero Knowledge Proof (ZKP). This, combined with optimizations for TLS 1.3, leads to an efficient protocol with constant communication in the online phase. Our work reduces online communication by 375× and online runtime by up to 4.6×, compared to prior work.

Open access
Scientific Computing and Data Management
Data Quality and Management
Research Data Management Practices
Original source
Mar 3, 2025·Proceedings of the ACM on software engineering.
0 cites
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair

Zaoyu Chen, Haoran Qin, Nuo Chen, Xiangyu Zhao · 7 authors

Smart contracts, predominantly written in Solidity and executed on blockchains like Ethereum, are immutable, making functional correctness paramount: once deployed, bugs and vulnerabilities become permanent. Despite rapid progress in transformer-based code LLMs, existing evaluations of Solidity code completion rely heavily on surface-form metrics (e.g., BLEU, CrystalBLEU) or hand-grading, which poorly correlate with functional correctness. Unlike Python, Solidity lacks large-scale and execution-based benchmarks, hindering systematic assessment and optimization of LLMs for smart contract development. To bridge this research gap, we introduce SolBench, a comprehensive benchmark and automated testing pipeline for Solidity, designed to emphasize functional correctness via differential fuzzing. SolBench contains 28,825 functions from 7,604 contracts collected from Etherscan (genesis to 2024), spanning 10 popular domains. We benchmark 14 diverse LLMs (open/closed, 1.3B to 671B parameters, general/code-specific, with/without reasoning). The dominant failure mode is missing crucial details (e.g., type definitions, state variables) in intra-contract context. Providing full-contract context mitigates this and improves code completion accuracy. However, full-context inference can be prohibitively expensive in practice. Generating outputs with large context windows using state-of-the-art models often incurs significant costs, rendering naive context scaling economically impractical. Crucially, most of a contract is irrelevant to implementing a given function; only a small subset of details is needed. To exploit this, we propose Retrieval-Augmented Repair (RAR), which integrates retrieval into code repair: it uses the executor's error messages to extract only the most relevant snippets from the full contract. RAR sharply reduces input length for function completion, improving accuracy while significantly cutting computational cost. We further analyze retrieval and code repair strategies within RAR, showing substantial improvements in accuracy and efficiency. SolBench and our RAR framework enable principled evaluation and cost-effective improvement of Solidity code generation. Dataset and code are available at https://github.com/ZaoyuChen/SolBench.

Open access
2 source records
cs.SE
cs.AI
cs.CL
Original source
Feb 6, 2025·arXiv (Cornell University)
1 cites
ExpProof : Operationalizing Explanations for Confidential Models with ZKPs

Chhavi Yadav, Evan Monroe Laufer, Dan Boneh, Kamalika Chaudhuri

In principle, explanations are intended as a way to increase trust in machine learning models and are often obligated by regulations. However, many circumstances where these are demanded are adversarial in nature, meaning the involved parties have misaligned interests and are incentivized to manipulate explanations for their purpose. As a result, explainability methods fail to be operational in such settings despite the demand \cite{bordt2022post}. In this paper, we take a step towards operationalizing explanations in adversarial scenarios with Zero-Knowledge Proofs (ZKPs), a cryptographic primitive. Specifically we explore ZKP-amenable versions of the popular explainability algorithm LIME and evaluate their performance on Neural Networks and Random Forests. Our code is publicly available at https://github.com/emlaufer/ExpProof.

Open access
2 source records
cs.LG
cs.AI
cs.CR
Original source
Jan 21, 2025·arXiv (Cornell University)
0 cites
Multi-source Multi-level Multi-token Ethereum Dataset and Benchmark Platform

Haoyuan Li, Mengxiao Zhang, Maoyuan Li, Jianzheng Li · 8 authors

This paper introduces 3MEthTaskforce (https://3meth.github.io), a multi-source, multi-level, and multi-token Ethereum dataset addressing the limitations of single-source datasets. Integrating over 300 million transaction records, 3,880 token profiles, global market indicators, and Reddit sentiment data from 2014-2024, it enables comprehensive studies on user behavior, market sentiment, and token performance. 3MEthTaskforce defines benchmarks for user behavior prediction and token price prediction tasks, using 6 dynamic graph networks and 19 time-series models to evaluate performance. Its multimodal design supports risk analysis and market fluctuation modeling, providing a valuable resource for advancing blockchain analytics and decentralized finance research.

Open access
2 source records
cs.CE
Scientific Computing and Data Management
Advanced Data Storage Technologies
Original source
Jan 1, 2025·IEEE Access
2 cites
Transchain: Blockchain-Based Management of Allografts for Enhancing Data Provenance

Rashid Ul Haq, Rahim Khan, Fahad Alturise, Shafrida Sahrani · 6 authors

Recent technological advances have enabled researchers to investigate various novel approaches utilized to manage allograft transplants and overcome the challenges of conventional centralized systems. The rising need for transparency, efficiency, and, especially, security in this highly sensitive medical procedure necessitates the use of decentralized solutions like blockchain rather than existing centralized approaches. However, the current state of research is theoretical and unproven, and allograft management lacks any reliable, cost-effective, or data-proven solution. In this paper, we propose an Ethereum blockchain-based allograft transplantation management system that can address all of those issues linked to the existing solutions. The proposed approach aims to enhance traceability, transparency, and data provenance across the entire allograft transplant process. We present six reliable and cost-efficient algorithms, as well as a comprehensive system architecture, to provide valuable insight into system implementation complexity. We have designed an efficient smart contract implementing the proposed algorithms to ensure flawless execution of allograft donation, transportation, and transplantation. We conduct thorough tests, validation, security, cost, throughput, and latency assessments of the system in order to contrast its effectiveness with existing solutions and results shows that our solution is cost-effective, as well as secure and efficient. We generalized the proposed solution so that, with minimal changes, it could be used for other problems and addressed some of the technical and ethical challenges.

Open access
Scientific Computing and Data Management
Research Data Management Practices
Data Quality and Management
Original source
Jan 1, 2025·IET Software
6 cites
Blockchain‐Audited Federated Learning: Securing Data and Model Updates With On‐Chain Provenance

Seid Mehammed, Girma Bewuketu, Demeke Getaneh, Md Nasre Alam · 6 authors

We present a permissioned blockchain–audited federated learning (FL) framework that strengthens data provenance and model‐update integrity. Our contribution is primarily engineering and architectural: a modular two‐channel design (provenance vs. update‐audit), lightweight on‐chain validation with off‐chain analytics, and a practical mapping to the 1 + 5 architectural views. In a TensorFlow Federated + Hyperledger Fabric prototype with 10 clients, we observe ≈18% faster anomaly detection under attack and a + 0.4 pp accuracy delta versus a baseline FL setup, with ~6% communication and ~8% energy overhead. We also provide a proof‐of‐concept zero‐knowledge succinct noninteractive argument of knowledge (zk‐SNARK) flow to validate per‐client summary properties off‐chain while anchoring results on‐chain. These contributions collectively advance the practical deployment of secure, auditable FL systems.

Open access
Scientific Computing and Data Management
Privacy-Preserving Technologies in Data
Blockchain Technology Applications and Security
Original source
Jan 1, 2025·Open MIND
0 cites
Architecting Autonomous Data Platforms: Integrating AI-Driven Governance, Metadata Intelligence, And Data Mesh Principles

Srinivasa Rao Seetala

Modern enterprises generate vast volumes of data across distributed applications, cloud platforms, and digital services. Traditional centralized data governance models struggle to scale in such complex environments, leading to data silos, inconsistent governance enforcement, and limited data accessibility. Autonomous data platforms supported by artificial intelligence (AI) offer a promising solution by integrating self-service infrastructure, automated governance mechanisms, and intelligent metadata management. AI-driven governance frameworks can automate tasks such as data discovery, classification, lineage tracking, anomaly detection, and compliance monitoring. This article explores the architectural foundations of autonomous data platforms and examines how AI-driven governance enables scalable, decentralized, and trustworthy data ecosystems. Drawing on emerging concepts such as data mesh architectures, federated governance models, and responsible AI frameworks, the paper proposes a conceptual model for building intelligent and self-governing enterprise data platforms. In such environments, machine learning algorithms continuously analyze data flows, schema evolution, usage patterns, and policy compliance to dynamically enforce governance rules and improve data quality. Metadata-driven architectures further enable automated cataloging, semantic enrichment, and real-time lineage tracking, allowing organizations to maintain transparency and accountability across complex data pipelines. By embedding governance directly into the data infrastructure, autonomous platforms reduce operational overhead while empowering domain teams to manage their own data products within standardized governance policies. Furthermore, the integration of explainable AI techniques and policy-aware automation ensures that governance decisions remain auditable, fair, and aligned with regulatory requirements. Ultimately, the convergence of AI, distributed data architectures, and intelligent metadata management provides a scalable foundation for building resilient, adaptive, and trustworthy enterprise data ecosystems capable of supporting advanced analytics, machine learning, and data-driven decision-making.

Open access
2 source records
Scientific Computing and Data Management
Data Quality and Management
Research Data Management Practices
Original source
Jan 1, 2025·Irish Interdisciplinary Journal of Science & Research
0 cites
Blockchain-Enabled Federated Learning Framework for Secure and Collaborative Drug Discovery: Integrating AI, Molecular Docking, and Distributed Ledger Technology

Mohamed Alshalaan, Nayyar Ahmed Khan

Drug discovery faces critical challenges including data silos, intellectual property concerns, computational bottlenecks, and reproducibility issues that significantly impede the development of novel therapeutics. This research proposes a novel Blockchain-enabled Federated Learning Framework for Drug Discovery (BFLD) that integrates distributed ledger technology, federated machine learning, and molecular docking simulations to create a secure, transparent, and collaborative ecosystem for pharmaceutical research. Our framework addresses key limitations in traditional drug discovery pipelines by enabling multi-institutional collaboration without compromising proprietary data, ensuring immutable audit trails for compound screening results, and accelerating hit-to-lead optimization through decentralized computing. We evaluate BFLD using datasets from 12 pharmaceutical research institutions, encompassing 2.4 million molecular compounds and 847 protein targets. Results demonstrate a 68% reduction in lead compound identification time, 91% improvement in data provenance tracking, and 94% stakeholder confidence in intellectual property protection. The framework achieves 89.7% accuracy in toxicity prediction through federated learning models while maintaining complete data privacy. Smart contracts automate licensing agreements and ensure equitable attribution of discoveries across participating institutions. This research establishes a paradigm shift toward decentralized, trustless pharmaceutical innovation aligned with open science principles while protecting commercial interests.

Open access
Blockchain Technology Applications and Security
Computational Drug Discovery Methods
Scientific Computing and Data Management
Original source
Jan 1, 2025·Procedia Computer Science
6 cites
NFT-based Data Provenance for AI Transparency in Enterprise Information Systems

Yiannis Verginadis, Orestis Almpanoudis, Dimitris Apostolou, Marcela Tuler de Oliveira · 5 authors

Enterprise Information Systems have a long-established and crucial role for modern organizations, as they enable seamless integration and management of critical business processes, ensuring efficiency in operations, data accuracy, and enhanced decision-making capabilities. One of their most interesting emerging technologies refer to the use of Artificial Intelligence as they may seamlessly automate routine tasks, offer predictive analytics, and provide deep insights, ultimately leading to intelligent data-driven decisions and improved operational efficiency. Of course, this direction of work is accompanied by some important challenges that come from the opacity of certain AI models and their potential biases due to low-quality training data used. In this paper, we argue that such challenges can be mitigated by a novel framework able to integrate, in a transparent manner, quality-related metadata on datasets used for training the AI-enabled emerging technologies in the field of EIS systems. These metadata are minted as Non-Fungible Tokens (NFTs) over the blockchain.

Open access
Scientific Computing and Data Management
Big Data and Business Intelligence
Ethics and Social Impacts of AI
Original source