Blockchain Papers

Follow blockchain research across journals, conferences, and preprint repositories.

122 papersLast indexed Aug 31, 2026
Search papers

Paper index

122 results · page 2 of 6

Clear filters
Jun 26, 2025·arXiv (Cornell University)
0 cites
ZKPROV: A Zero-Knowledge Approach to Dataset Provenance for Large Language Models

Mina Namazi, Alexander Nemecek, Erman Ayday

As large language models (LLMs) are used in sensitive fields, accurately verifying their computational provenance without disclosing their training datasets poses a significant challenge, particularly in regulated sectors such as healthcare, which have strict requirements for dataset use. Traditional approaches either incur substantial computational cost to fully verify the entire training process or leak unauthorized information to the verifier. Therefore, we introduce ZKPROV, a novel cryptographic framework allowing users to verify that the LLM's responses to their prompts are trained on datasets certified by the authorities that own them. Additionally, it ensures that the dataset's content is relevant to the users' queries without revealing sensitive information about the datasets or the model parameters. ZKPROV offers a unique balance between privacy and efficiency by binding training datasets, model parameters, and responses, while also attaching zero-knowledge proofs to the responses generated by the LLM to validate these claims. Our experimental results demonstrate sublinear scaling for generating and verifying these proofs, with end-to-end overhead under 3.3 seconds for models up to 8B parameters, presenting a practical solution for real-world applications. We also provide formal security guarantees, proving that our approach preserves dataset confidentiality while ensuring trustworthy dataset provenance.

Open access
2 source records
cs.CR
cs.AI
cs.LG
Original source
Jun 17, 2025·arXiv (Cornell University)
0 cites
Explain First, Trust Later: LLM-Augmented Explanations for Graph-Based Crypto Anomaly Detection

Watson, Adriana, Richards, Grant, Schiff, Daniel

The decentralized finance (DeFi) community has grown rapidly in recent years, pushed forward by cryptocurrency enthusiasts interested in the vast untapped potential of new markets. The surge in popularity of cryptocurrency has ushered in a new era of financial crime. Unfortunately, the novelty of the technology makes the task of catching and prosecuting offenders particularly challenging. Thus, it is necessary to implement automated detection tools related to policies to address the growing criminality in the cryptocurrency realm.

Open access
2 source records
cs.CE
cs.AI
cs.CR
Original source
Jun 9, 2025·Anais Estendidos do XXV Simpósio Brasileiro de Computação Aplicada à Saúde (SBCAS 2025)
0 cites
MEPCA: a technical model to improve on-chain Electronic Health Records processing

Fausto Neri da Silva Vanin, Rodrigo da Rosa Righi, Cristiano André da Costa

Blockchain technology in healthcare is gaining attention for addressing data privacy, interoperability, and health record integrity issues. Standards like HL7 FHIR and OpenEHR ensure data consistency, but privacy concerns persist under regulations like HIPAA, GDPR, and LGPD. Existing methods often store only data hashes, raising validation risks. The MEPCA model introduces a blockchain-based framework for secure health record management, focusing on on-chain EHR data processing. Key elements include Data Steward, Shared Data Vault, and Zero-Knowledge Proofs of HL7 FHIR fields. Experiments with Fully Homomorphic Encryption show enhanced security and reliability for health records, offering a robust alternative to traditional off-chain approaches.

Open access
Data Quality and Management
Electronic Health Records Systems
Artificial Intelligence in Healthcare
Original source
May 20, 2025·arXiv (Cornell University)
0 cites
Zk-SNARK for String Match

T. Li, Liao, Taobo

We present a secure and efficient string-matching platform leveraging zk-SNARKs (Zero-Knowledge Succinct Non-Interactive Arguments of Knowledge) to address the challenge of detecting sensitive information leakage while preserving data privacy. Our solution enables organizations to verify whether private strings appear on public platforms without disclosing the strings themselves. To achieve computational efficiency, we integrate a sliding window technique with the Rabin-Karp algorithm and Rabin Fingerprint, enabling hash-based rolling comparisons to detect string matches. This approach significantly reduces time complexity compared to traditional character-by-character comparisons. We implement the proposed system using gnark, a high-performance zk-SNARK library, which generates succinct and verifiable proofs for privacy-preserving string matching. Experimental results demonstrate that our solution achieves strong privacy guarantees while maintaining computational efficiency and scalability. This work highlights the practical applications of zero-knowledge proofs in secure data verification and contributes a scalable method for privacy-preserving string matching.

Open access
2 source records
cs.CR
Data Quality and Management
Web Application Security Vulnerabilities
Original source
Apr 3, 2025·Scientific Journal of Artificial Intelligence and Blockchain Technologies
0 cites
Healthcare Data Provenance Using Distributed Ledger Systems

Niharika Singh

The unprecedented growth of digital health ecosystems, fueled by electronic health records (EHRs), wearable devices, telemedicine, and AI-driven diagnostics, has amplified the critical need for reliable data provenance mechanisms. Provenance, defined as the comprehensive history of data generation, access, transformation, and transfer, ensures that stakeholders—including patients, clinicians, insurers, researchers, and regulators—can trust the authenticity, integrity, and accountability of healthcare information. Traditional provenance systems, often centralized, are vulnerable to insider manipulation, cyberattacks, data silos, and audit inefficiencies, thereby undermining trust and regulatory compliance. Distributed Ledger Systems (DLS), encompassing blockchain, permissioned ledgers, and Directed Acyclic Graphs (DAGs), offer a paradigm shift by enabling immutable, transparent, and tamper-evident provenance trails across diverse healthcare stakeholders. This manuscript provides an in-depth exploration of DLS-enabled healthcare data provenance by reviewing current literature, identifying research gaps, and developing a methodological framework tested through simulated experiments. Empirical evaluation demonstrates that distributed ledgers reduce provenance validation time by 57–71%, accelerate audit processes by up to 70%, and significantly enhance regulatory traceability under HIPAA and GDPR requirements. Moreover, patient-centric smart contracts and decentralized identifiers foster individual ownership and interoperability, reshaping data governance models toward inclusivity and transparency. While challenges such as scalability, energy efficiency, and privacy-preserving erasure remain, the findings highlight DLS as a transformative infrastructure for establishing trustworthy healthcare ecosystems. The study concludes by recommending hybrid ledger architectures, cryptographic privacy enhancements, and supportive policy frameworks to ensure sustainable, ethical, and globally interoperable healthcare data provenance systems.

Open access
Scientific Computing and Data Management
Blockchain Technology Applications and Security
Data Quality and Management
Original source
Mar 26, 2025·arXiv
2 cites
A Blockchain-Enabled Framework for Storage and Retrieval of Social Data

Aishwarya Parab, P. Pradhan, Yogesh Simmhan, Arnab K. Paul

The increasing availability of data from diverse sources, including trusted entities such as governments, as well as untrusted crowd-sourced contributors, demands a secure and trustworthy environment for storage and retrieval. Blockchain, as a distributed and immutable ledger, offers a promising solution to address these challenges. This short paper studies the feasibility of a blockchain-based framework for secure data storage and retrieval across trusted and untrusted sources, focusing on provenance, storage mechanisms, and smart contract security. Through initial experiments using Hyper Ledger Fabric (HLF), we evaluate the storage efficiency, scalability, and feasibility of the proposed approach. This study serves as a motivation for future research to develop a comprehensive blockchain-based storage and retrieval framework.

Open access
2 source records
cs.DC
Blockchain Technology Applications and Security
Data Quality and Management
Original source
Mar 26, 2025·Symmetry
3 cites
Enhancing Cryptocurrency Security: Leveraging Embeddings and Large Language Models for Creating Cryptocurrency Security Expert Systems

Ahmed Mohamed Abdallah, Heba K. Aslan, Mohamed S. Abdallah, Young Im Cho · 5 authors

In recent years, the rapid growth of cryptocurrency markets has highlighted the urgent need for advanced security solutions capable of addressing a spectrum of unique threats, from phishing and wallet hacks to complex blockchain vulnerabilities. This paper presents a comprehensive approach to fortifying cryptocurrency systems by harnessing the structural symmetry inherent in transactional patterns. By leveraging local large language models (LLMs), embeddings, and vector databases, we develop an intelligent and scalable security expert system that exploits symmetry-based anomaly detection to enhance threat identification. Cryptocurrency networks face increasing threats from sophisticated attacks that often exploit asymmetric vulnerabilities. To counteract these risks, we propose a novel security expert system that integrates symmetry-aware analysis through LLMs and advanced embedding techniques. Our system efficiently captures symmetrical transaction patterns, enabling robust detection of anomalies and threats while preserving structural integrity. By integrating a modular framework with LangChain and a vector database (Chroma DB), we achieve improved accuracy, recall, and precision by leveraging the symmetry of transaction distributions and behavioral patterns. This work sets a new benchmark for LLM-driven cybersecurity solutions, offering a scalable and adaptive approach to reinforcing the security symmetry in cryptocurrency systems. The proposed expert system was evaluated using a benchmark dataset of cryptocurrency transactions, including real-world threat scenarios involving phishing, fraudulent transactions, and blockchain anomalies. The system achieved an accuracy of 92%, a precision of 89%, and a recall of 93%, demonstrating a 10% improvement over existing security frameworks. Compared to traditional rule-based and machine learning-based detection methods, our approach significantly enhances real-time threat detection while reducing false positives. The integration of LLMs with embeddings and vector retrieval enables more efficient contextual anomaly detection, setting a new benchmark for AI-driven security solutions in the cryptocurrency domain.

Open access
Data Quality and Management
Blockchain Technology Applications and Security
Advanced Malware Detection Techniques
Original source
Mar 10, 2025·Brazilian Review of Finance
2 cites
Forecasting Bitcoin and Ethereum risk measures through MSGARCH models

Luiz Koodi Hotta, Carlos Trucíos, Pedro L. Valls Pereira, Mauricio Zevallos

Recent studies have suggested that more complex models than GARCH are better suited for forecasting cryptocurrency risk measures, such as Value-at-Risk and Expected Shortfall. Among these studies, some highlight the advantages of MSGARCH models over traditional GARCH models. While improvements over single-regime GARCH models have been observed by using MSGARCH, the literature has only focused on the MSGARCH specification proposed by Haas, Mittnik and Paolella (Journal of Financial Econometrics, 2004) overlooking several other well-established MSGARCH specification alternatives. In this paper, we illustrate that exploring alternative MSGARCH specifications can lead to improvements in risk measure performance, emphasizing the potential benefits of using several specifications.

Open access
Big Data Technologies and Applications
Data Quality and Management
Probability and Risk Models
Original source
Mar 7, 2025·Proceedings on Privacy Enhancing Technologies
1 cites
ORIGO: Proving Provenance of Sensitive Data with Constant Communication

Jens Ernstberger, Jan Lauinger, Yulin Wu, Arthur Gervais · 5 authors

Transport Layer Security (TLS) is foundational for safeguarding client-server communication. However, it does not extend integrity guarantees to third-party verification of data authenticity. If a client wants to present data obtained from a server, it cannot convince any other party that the data has not been tampered with. TLS oracles ensure data authenticity beyond the client-server TLS connection, such that clients can obtain data from a server and ensure provenance to any third party, without server-side modifications. Generally, a TLS oracle involves a third party, the verifier, in a TLS session to verify that the data obtained by the client is accurate. Existing protocols for TLS oracles are communication-heavy, as they rely on interactive protocols. We present ORIGO, a TLS oracle with constant communication. Similar to prior work, ORIGO introduces a third party in a TLS session, and provides a protocol to ensure the authenticity of data transmitted in a TLS session, without forfeiting its confidentiality. Compared to prior work, we rely on intricate details specific to TLS 1.3, which allow us to prove correct key derivation, authentication and encryption within a Zero Knowledge Proof (ZKP). This, combined with optimizations for TLS 1.3, leads to an efficient protocol with constant communication in the online phase. Our work reduces online communication by 375× and online runtime by up to 4.6×, compared to prior work.

Open access
Scientific Computing and Data Management
Data Quality and Management
Research Data Management Practices
Original source
Feb 7, 2025·arXiv (Cornell University)
4 cites
Mining a Decade of Event Impacts on Contributor Dynamics in Ethereum: A Longitudinal Study

Matteo Vaccargiu, Sabrina Aufiero, Cheikh Oumar Ba, Silvia Bartolucci · 9 authors

We analyze developer activity across 10 major Ethereum repositories (totaling 129884 commits, 40550 issues) spanning 10 years to examine how events such as technical upgrades, market events, and community decisions impact development. Through statistical, survival, and network analyses, we find that technical events prompt increased activity before the event, followed by reduced commit rates afterwards, whereas market events lead to more reactive development. Core infrastructure repositories like Go-Ethereum exhibit faster issue resolution compared to developer tools, and technical events enhance core team collaboration. Our findings show how different types of events shape development dynamics, offering insights for project managers and developers in maintaining development momentum through major transitions. This work contributes to understanding the resilience of development communities and their adaptation to ecosystem changes.

Open access
3 source records
Software Engineering Research
Software System Performance and Reliability
Data Quality and Management
Original source
Jan 13, 2025·IACR Communications in Cryptology
6 cites
Foundations of Data Availability Sampling

Mathias Hall-Andersen, Mark Simkin, Benedikt Wagner

Towards building more scalable blockchains, an approach known as data availability sampling (DAS) has emerged over the past few years. Even large blockchains like Ethereum are planning to eventually deploy DAS to improve their scalability. In a nutshell, DAS allows the participants of a network to ensure the full availability of some data without any one participant downloading it entirely. Despite the significant practical interest that DAS has received, there are currently no formal definitions for this primitive, no security notions, and no security proofs for any candidate constructions. For a cryptographic primitive that may end up being widely deployed in large real-world systems, this is a rather unsatisfactory state of affairs. In this work, we initiate a cryptographic study of data availability sampling. To this end, we define data availability sampling precisely as a clean cryptographic primitive. Then, we show how data availability sampling relates to erasure codes. We do so by defining a new type of commitment schemes which naturally generalizes vector commitments and polynomial commitments. Using our framework, we analyze existing constructions and prove them secure. In addition, we give new constructions which are based on weaker assumptions, computationally more efficient, and do not rely on a trusted setup, at the cost of slightly larger communication complexity. Finally, we evaluate the trade-offs of the different constructions.

Open access
Cloud Data Security Solutions
Data Quality and Management
Blockchain Technology Applications and Security
Original source
Jan 1, 2025·IEEE Access
2 cites
Transchain: Blockchain-Based Management of Allografts for Enhancing Data Provenance

Rashid Ul Haq, Rahim Khan, Fahad Alturise, Shafrida Sahrani · 6 authors

Recent technological advances have enabled researchers to investigate various novel approaches utilized to manage allograft transplants and overcome the challenges of conventional centralized systems. The rising need for transparency, efficiency, and, especially, security in this highly sensitive medical procedure necessitates the use of decentralized solutions like blockchain rather than existing centralized approaches. However, the current state of research is theoretical and unproven, and allograft management lacks any reliable, cost-effective, or data-proven solution. In this paper, we propose an Ethereum blockchain-based allograft transplantation management system that can address all of those issues linked to the existing solutions. The proposed approach aims to enhance traceability, transparency, and data provenance across the entire allograft transplant process. We present six reliable and cost-efficient algorithms, as well as a comprehensive system architecture, to provide valuable insight into system implementation complexity. We have designed an efficient smart contract implementing the proposed algorithms to ensure flawless execution of allograft donation, transportation, and transplantation. We conduct thorough tests, validation, security, cost, throughput, and latency assessments of the system in order to contrast its effectiveness with existing solutions and results shows that our solution is cost-effective, as well as secure and efficient. We generalized the proposed solution so that, with minimal changes, it could be used for other problems and addressed some of the technical and ethical challenges.

Open access
Scientific Computing and Data Management
Research Data Management Practices
Data Quality and Management
Original source
Jan 1, 2025·Data & Policy
2 cites
Data technologies and analytics for policy and governance: a landscape review

Omar Isaac Asensio, Catherine E. Moore, Nícola Ulibarrí, Mecit Can Emre Simsekler · 6 authors

Abstract Data for Policy ( dataforpolicy.org ), a trans-disciplinary community of research and practice, has emerged around the application and evaluation of data technologies and analytics for policy and governance. Research in this area has involved cross-sector collaborations, but the areas of emphasis have previously been unclear. Within the Data for Policy framework of six focus areas, this report offers a landscape review of Focus Area 2: Technologies and Analytics. Taking stock of recent advancements and challenges can help shape research priorities for this community. We highlight four commonly used technologies for prediction and inference that leverage datasets from the digital environment: machine learning (ML) and artificial intelligence systems, the internet-of-things, digital twins, and distributed ledger systems. We review innovations in research evaluation and discuss future directions for policy decision-making.

Open access
Data Quality and Management
Big Data and Business Intelligence
Privacy-Preserving Technologies in Data
Original source
Jan 1, 2025·Open MIND
0 cites
Architecting Autonomous Data Platforms: Integrating AI-Driven Governance, Metadata Intelligence, And Data Mesh Principles

Srinivasa Rao Seetala

Modern enterprises generate vast volumes of data across distributed applications, cloud platforms, and digital services. Traditional centralized data governance models struggle to scale in such complex environments, leading to data silos, inconsistent governance enforcement, and limited data accessibility. Autonomous data platforms supported by artificial intelligence (AI) offer a promising solution by integrating self-service infrastructure, automated governance mechanisms, and intelligent metadata management. AI-driven governance frameworks can automate tasks such as data discovery, classification, lineage tracking, anomaly detection, and compliance monitoring. This article explores the architectural foundations of autonomous data platforms and examines how AI-driven governance enables scalable, decentralized, and trustworthy data ecosystems. Drawing on emerging concepts such as data mesh architectures, federated governance models, and responsible AI frameworks, the paper proposes a conceptual model for building intelligent and self-governing enterprise data platforms. In such environments, machine learning algorithms continuously analyze data flows, schema evolution, usage patterns, and policy compliance to dynamically enforce governance rules and improve data quality. Metadata-driven architectures further enable automated cataloging, semantic enrichment, and real-time lineage tracking, allowing organizations to maintain transparency and accountability across complex data pipelines. By embedding governance directly into the data infrastructure, autonomous platforms reduce operational overhead while empowering domain teams to manage their own data products within standardized governance policies. Furthermore, the integration of explainable AI techniques and policy-aware automation ensures that governance decisions remain auditable, fair, and aligned with regulatory requirements. Ultimately, the convergence of AI, distributed data architectures, and intelligent metadata management provides a scalable foundation for building resilient, adaptive, and trustworthy enterprise data ecosystems capable of supporting advanced analytics, machine learning, and data-driven decision-making.

Open access
2 source records
Scientific Computing and Data Management
Data Quality and Management
Research Data Management Practices
Original source
Dec 31, 2024·Empirical Software Engineering
4 cites
UPC sentinel: An accurate approach for detecting upgradeability proxy contracts in Ethereum

Amir M. Ebrahimi, Bram Adams, Gustavo A. Oliva, Ahmed E. Hassan

Software applications that run on a blockchain platform are known as DApps. DApps are built using smart contracts, which are immutable after deployment. Just like any real-world software system, DApps need to receive new features and bug fixes over time in order to remain useful and secure. However, Ethereum lacks native solutions for post-deployment smart contract maintenance, requiring developers to devise their own methods. A popular method is known as the upgradeability proxy contract (UPC), which involves implementing the proxy design pattern (as defined by the Gang of Four). In this method, client calls first hit a proxy contract, which then delegates calls to a certain implementation contract. Most importantly, the proxy contract can be reconfigured during runtime to delegate calls to another implementation contract, effectively enabling application upgrades. For researchers, the accurate detection of UPCs is a strong requirement in the understanding of how exactly real-world DApps are maintained over time. For practitioners, the accurate detection of UPCs is crucial for providing application behavior transparency and enabling auditing. In this paper, we introduce UPC Sentinel, a novel three-layer algorithm that utilizes both static and dynamic analysis of smart contract bytecode to accurately detect active UPCs. We evaluated UPC Sentinel using two distinct ground truth datasets. In the first dataset, our method demonstrated a near-perfect accuracy of 99%. The evaluation on the second dataset further established our method's efficacy, showing a perfect precision rate of 100% and a near-perfect recall of 99.3%, outperforming the state of the art. Finally, we discuss the potential value of UPC Sentinel in advancing future research efforts.

Open access
4 source records
Software Engineering Research
Data Quality and Management
Imbalanced Data Classification Techniques
Original source
Dec 28, 2024·Scientific Reports
3 cites
An improved practical Byzantine fault tolerance algorithm for aggregating node preferences

Xu Liu, Junwu Zhu

Consensus algorithms play a critical role in maintaining the consistency of blockchain data, directly affecting the system's security and stability, and are used to determine the binary consensus of whether proposals are correct. With the development of blockchain-related technologies, social choice issues such as Bitcoin scaling and main chain forks, as well as the proliferation of decentralized autonomous organization (DAO) applications based on blockchain technology, require consensus algorithms to reach consensus on a specific proposal among multiple proposals based on node preferences, thereby addressing the multi-value consensus problem. However, existing consensus algorithms, including Practical Byzantine Fault Tolerance (PBFT), do not support nodes expressing preferences. Instead, the proposal to reach consensus is directly decided by specific nodes, with other nodes merely verifying the proposal's validity, which can easily result in monopolistic or dictatorial outcomes. In response, we proposed the Aggregating Preferences with Practical Byzantine Fault Tolerance (AP-PBFT) consensus algorithm, which allows nodes to express preferences for multiple proposals. AP-PBFT ensures the validity of consensus results through a consensus output protocol, and incentivizes nodes to act honestly during the consensus process by incentive mechanism. First, AP-PBFT leverages Verifiable Random Function to select both consensus nodes and a primary node from the candidates. The primary node gathers proposals, assembles them into a proposal package, and broadcasts it to other consensus nodes. The consensus nodes independently vote to express their preferences for different proposals in the package, execute the consensus output protocol to reach local consensus, and the primary node aggregates these results to form the global consensus. Once the global consensus is finalized, AP-PBFT evaluates node behavior based on the consensus output protocol, penalizes nodes that acted maliciously, and rewards those that adhered to the protocol. Additionally, nodes can interact and adopt different strategies while executing the consensus output protocol, which can influence the consensus outcome. Therefore, we established an evolutionary game model based on hypergraph to analyze these interactions. Theoretical analysis shows that the incentive mechanism in AP-PBFT effectively encourages nodes to honestly follow the consensus output protocol, ensuring that AP-PBFT satisfies the properties of consistency, validity, and termination. Finally, the simulation results demonstrate that the AP-PBFT algorithm possesses good scalability and the capability to handle dynamic changes in nodes, surpassing some mainstream consensus algorithms in terms of transaction throughput and consensus achievement time. Moreover, AP-PBFT can incentivize honest behavior among consensus nodes, thereby enhancing the reliability of consensus and strengthening the security of the network.

Open access
Distributed systems and fault tolerance
Cloud Computing and Resource Management
Data Quality and Management
Original source
Dec 27, 2024·IEEE Transactions on Information Forensics and Security
18 cites
Blockchain-Empowered Keyword Searchable Provable Data Possession for Large Similar Data

Ying Miao, Keke Gai, Jing Yu, Yu‐an Tan · 6 authors

Provable Data Possession (PDP) is an alternative technique that guarantees the integrity of remote data. However, most current PDP schemes are inapplicable to similarity-like data checking with the same attribute, i.e., when there are numerous similar files to be checked by Data Owners (DOs). Some traditional models cannot resist the corrupt auditors who always generate biased challenge information. Besides, a copy-summation attack exists in some schemes, which means the Cloud Server (CS) can bypass the verification by storing the median value instead of initial data via summation operation. To address the issues above, in this work, we propose a keyword searchable PDP scheme for large similar data checking. To achieve searchability, we introduce the notion of a keyword in PDP and design a specific index structure to match the authenticator. The scheme enables all matched files to be auditable and verifiable, while guaranteeing privacy protections. Unlike existing methods, our Third Party Auditor (TPA) checks all similar data containing the same keyword simultaneously. We utilize unpredictable yet verifiable public information on the blockchain to generate challenge information, rather than relying on a centralized TPA. The proposed scheme can resist copy-summation attacks. Theoretical analysis demonstrates that the proposed scheme satisfies the security requirements, and our evaluations demonstrate its efficiency.

Open access
Data Quality and Management
Original source
Dec 4, 2024·Proceedings of the International Conference on AI Research.
2 cites
LLM Supply Chain Provenance: A Blockchain-based Approach

Shridhar R. Singh, Luke Vorster

The burgeoning size and complexity of Large Language Models (LLMs) introduce significant challenges in ensuring data integrity. The proliferation of "deep fakes" and manipulated information raises concerns about the vulnerability of LLMs to misinformation. Traditional LLM architectures often lack robust mechanisms for tracking the origin and history of training data. This opaqueness can leave LLMs susceptible to manipulation by malicious actors who inject biased or inaccurate data. This research proposes a novel approach integrating Blockchain Technology (BCT) within the LLM data supply chain. With its core principle of a distributed and immutable ledger, BCT offers a compelling solution to address this challenge. By storing the LLM's data supply chain on a blockchain, we establish a verifiable record of data provenance. This allows for tracing the origin of each data point used to train the LLM, fostering greater transparency and trust in the model's outputs. This decentralised approach minimises the risk of single points of failure and manipulation. Additionally, the immutability of blockchain records ensures that the data provenance remains tamper-proof, further enhancing the trustworthiness of the LLM. Our approach leverages three critical features of BCT to strengthen LLM security: 1) Transaction Anonymity: While data provenance is recorded on the blockchain, identities of data contributors can be anonymised, protecting their privacy while ensuring data integrity. 2) Decentralised Repository: Enhances the system's resilience against potential attacks by distributing the data provenance record across the blockchain network. 3) Block Validation: Rigorous consensus mechanisms ensure the validity of each data point added to the LLM's data supply chain - minimising the risk of incorporating inaccurate or manipulated data into the training process. Using the experimental approach, initial evaluations using simulated LLM training data on a blockchain platform demonstrate the feasibility and effectiveness of the proposed approach in enhancing data integrity. This approach has far-reaching implications for ensuring the trustworthiness of LLMs in various applications.

Open access
Blockchain Technology Applications and Security
Data Quality and Management
Cloud Data Security Solutions
Original source
Nov 28, 2024·arXiv (Cornell University)
2 cites
Know Your Account: Double Graph Inference-Based Account De-Anonymization on Ethereum

Shuyi Miao, Wangjie Qiu, Hongwei Zheng, Qinnan Zhang · 9 authors

The scaled Web 3.0 digital economy, represented by decentralized finance (DeFi), has sparked increasing interest in the past few years, which usually relies on blockchain for token transfer and diverse transaction logic. However, illegal behaviors, such as financial fraud, hacker attacks, and money laundering, are rampant in the blockchain ecosystem and seriously threaten its integrity and security. In this paper, we propose a novel double graph-based Ethereum account de-anonymization inference method, dubbed DBG4ETH, which aims to capture the behavioral patterns of accounts comprehensively and has more robust analytical and judgment capabilities for current complex and continuously generated transaction behaviors. Specifically, we first construct a global static graph to build complex interactions between the various account nodes for all transaction data. Then, we also construct a local dynamic graph to learn about the gradual evolution of transactions over different periods. Different graphs focus on information from different perspectives, and features of global and local, static and dynamic transaction graphs are available through DBG4ETH. In addition, we propose an adaptive confidence calibration method to predict the results by feeding the calibrated weighted prediction values into the classifier. Experimental results show that DBG4ETH achieves state-of-the-art results in the account identification task, improving the F1-score by at least 3.75% and up to 40.52% compared to processing each graph type individually and outperforming similar account identity inference methods by 5.23 % to 12.91 %.

Open access
3 source records
Data Quality and Management
Privacy-Preserving Technologies in Data
Cloud Data Security Solutions
Original source
Nov 10, 2024·Proceedings on Privacy Enhancing Technologies
9 cites
Janus: Fast Privacy-Preserving Data Provenance For TLS

Jan Lauinger, Jens Ernstberger, Andreas Finkenzeller, Sebastian Steinhorst

Web users can gather data from secure endpoints and demonstrate the provenance of sensitive data to any third party by using privacy-preserving TLS oracles. In practice, privacy-preserving TLS oracles remain limited and cannot verify larger, sensitive data sets. In this work, we introduce new optimizations for TLS oracles, which enhance the efficiency of selectively verifying the provenance of confidential web data. The novelty of our work is a construction which secures an honest verifier zero-knowledge proof system in the asymmetric privacy setting while retaining security against malicious adversaries. Concerning TLS 1.3 in the one round-trip time (1-RTT) mode, we propose a new, optimized garble-then-prove paradigm in a security setting with malicious adversaries. Our improvements reach new performance benchmarks and facilitate a practical deployment of privacy-preserving TLS oracles in web browsers.

Open access
Scientific Computing and Data Management
Data Quality and Management
Cloud Data Security Solutions
Original source
Oct 14, 2024·arXiv (Cornell University)
0 cites
Mastering AI: Big Data, Deep Learning, and the Evolution of Large Language Models -- Blockchain and Applications

Pohsun Feng, Ziqian Bi, Yan, Lawrence K. Q., Yizhu Wen · 17 authors

A detailed exploration of blockchain technology and its applications across various fields is provided, beginning with an introduction to cryptography fundamentals, including symmetric and asymmetric encryption, and their roles in ensuring security and trust within blockchain systems. The structure and mechanics of Bitcoin and Ethereum are then examined, covering topics such as proof-of-work, proof-of-stake, and smart contracts. Practical applications of blockchain in industries like decentralized finance (DeFi), supply chain management, and identity authentication are highlighted. The discussion also extends to consensus mechanisms and scalability challenges in blockchain, offering insights into emerging technologies like Layer 2 solutions and cross-chain interoperability. The current state of academic research on blockchain and its potential future developments are also addressed.

Open access
2 source records
Big Data and Digital Economy
Big Data Technologies and Applications
Data Quality and Management
Original source
Sep 26, 2024·World Journal of Advanced Research and Reviews
0 cites
Innovative approaches in data management and cybersecurity: Insights from recent studies

Mikhalov Alojo

The increasing complexity of data management systems, coupled with the evolving nature of cybersecurity threats, necessitates innovative approaches to ensure data integrity, confidentiality, and availability. This paper explores recent studies on advanced data management strategies and their intersection with cybersecurity practices. Key insights are drawn from the latest research on topics such as distributed ledger technologies, artificial intelligence-driven threat detection, and privacy-preserving data management frameworks. The analysis highlights how these emerging technologies are reshaping the landscape of data management while addressing cybersecurity challenges. Additionally, this paper examines the role of regulation and policy in fostering secure data ecosystems. The findings offer a comprehensive overview of current trends, challenges, and opportunities in the field, with recommendations for future research directions.

Open access
Big Data and Business Intelligence
Information and Cyber Security
Data Quality and Management
Original source
Sep 25, 2024·WSEAS Transactions on Information Science and Applications archive
7 cites
Enhancing the Reliability of Academic Document Certification Systems with Blockchain and Large Language Models

Jean Gilbert Mbula Mboma, Obed Tshimanga Tshipata, Witesyavwirwa Vianney Kambale, Mohamed Salem · 6 authors

Verifying the authenticity of documents, whether digital or physical, is a complex and crucial challenge faced by a variety of entities, including governments, regulators, financial institutions, educational establishments, and healthcare services. Rapid advances in technology have facilitated the creation of falsified or fraudulent documents, calling into question the credibility and authenticity of academic records. Most existing blockchain-based verification methods and systems focus primarily on verifying the integrity of a document, paying less attention to examining the authenticity of the document’s actual content before it is validated and registered in the system, thus opening loopholes for clever forgeries or falsifications. This paper details the design and implementation of a proof-of-concept system that combines GPT-3.5’s natural language processing prowess with the Ethereum blockchain and the InterPlanetary File System (IPFS) for storing and verifying documents. It explains how a Large Language Model like GPT-3.5 extracts essential information from academic documents and encrypts it before storing it in the blockchain ensuring document integrity and authenticity. The system is tested for its efficiency in handling both digital and physical documents, demonstrating increased security and reliability in academic document verification.

Open access
Blockchain Technology Applications and Security
Cloud Data Security Solutions
Data Quality and Management
Original source