The state-of-the-art for auditing and reproducing scientific applications on high-performance computing (HPC) systems is through a data provenance subsystem. While recent advances in data provenance lie in reducing the performance overhead and improving the user's query flexibility, the fidelity of data provenance is often overlooked: there is no such way to ensure that the provenance data itself has not been fabricated or falsified. This paper advocates leveraging blockchains to deliver immutable and autonomous data provenance services such that scientific discoveries are trustworthy. The challenges for adopting blockchains to HPC include designing a new blockchain architecture compatible with the HPC platforms and, more importantly, a set of new consensus protocols for scientific applications atop blockchains. To this end, we have designed the proof-of-scalable-traceability (POST) protocol and implemented it in a blockchain prototype, namely SciChain, the very first practical blockchain system for provenance services on HPC. We evaluated SciChain by comparing it with multiple state-of-the-art systems; experimental results showed that SciChain guaranteed trustworthy data provenance while incurring orders of magnitude lower overhead than existing solutions.
Scientific Computing and Data Management
Innovative Microfluidic and Catalytic Techniques Innovation
Zhoujie Zhang, Can Cui, Lei Tao, Jiaqi Wang ¡ 12 authors
In the era of big data and artificial intelligence development, data has become an important asset. Open data ecology has become an inevitable choice to promote business success. In the process of data open sharing, data provenance technology can supervise the data sharing path among multiple entities, and protect the interests of data owners. This paper proposes a data provenance technology for dispatching and control data based on blockchain. Firstly, a data provenance model is designed based on blockchain, which describes the provenance process of dispatching and control model data.Secondly, the distributed ledger of blockchain is used to record the data sharing path described by the data traceability model.Finally, the automatic implementation of provenance process is realized by using smart contract of blockchain.
In order to create a transparent and sound academic communication ecosystem centered on researchers, we developed a system that applied blockchain technology to an open peer review system. In this study, an open peer review system was developed based on Hyperledger Fabric, which is a private blockchain. The system can be operated in connection with the reviewer recommendation module of the existing submission management system. In the reviewer recommendation module, reviewers are recommended by excluding co-authors and colleagues after an expertise test. The blockchain system performs an open peer review process based on smart contracts, while the submission management system selects reviewers for peer review. A service broker intervenes between these two systems for data interchange. The system developed herein is expected to be used as a researcher-centered scholarly communication model in the open science era, in which the intervention of publishers is minimized, and authors and reviewers (as researchers) are centered.
In the past few decades, there has been a sharp rise of research irreproducibility and retraction, to a point that now is deemed as a crisis. Addressing this crisis, we present a peer-to-peer (P2P) publication model that utilizes blockchain and smart contract technologies. Focusing primarily on researchers and reviewers, the conceptual P2P publication model addresses the sociocultural and incentivization aspects of the irreproducibility crisis. In the P2P publication model, instead of a complete publication, a preapproved experimental design will be published on an incremental basis (unit-by-unit) and authorship will be shared with reviewers. The concept of the P2P publication model was inspired by the transformational journey the music publishing industry has undertaken as it traverses through vinyl age (complete albums) to the Spotify age (single-by-single), where there is a growing inclination among artists toward building an incremental album, taking account of feedback from fans and utilizing automated revenue collection and sharing systems. The ability to publish incrementally through the P2P publication model will relieve researchers from the burden of publishing complete and âgood resultsâ while simultaneously incentivizing reviewers to undertake rigorous review work to gain authorship credit in the research. The proposed P2P publication model aims to transform the century-old publication model and incentivization structure in alignment with open access publication ethos of the 21st century.
This personal reaction is written from multiple perspectives. First and foremost, as the corresponding author of the original FAIR article. Second as the chair of the first High Level Expert Group (HLEG) of European Open Science Cloud (EOSC) (which is how I met Jean-Claude) and third from my current GO FAIR and CODATA perspective. None of what I write below is to be seen as a formal position of any of the organisations I am associated with.Let me start by stating that, after some periods silent of hope and of deep despair, I now strongly feel that, with the governance of the EOSC Association in place, EOSC will become a success after all. It will still be critical that the Association involves the member states (MSs) and actual researchers in an agile and non-bureaucratic manner, for which we need bottom-up mechanisms such as operated by the Research Data Alliance (RDA) and GO FAIR. But a balancing formal entity operating along the formalised Strategic Research and Innovation Agenda [1] and the Partnership proposal as well as the various âdeclarationsâ including the recent one under the German presidency [2] are an excellent guiding roadmap to a successful EOSC, obviously in global context.That said, at the risk of sounding like broken record, this reaction should also look at the points where it went âalmostâ wrong, as we should try and learn from our mistakes. I may make some enemiesâor strengthen the opinion of existing onesâin the process, but then, a wise old friend, who also wrote one of the reactions once told me: âBarend, unless you made some enemies you probably lived in vain.â So I will speak my mind (âwhat's new'?). I also like to say that âEOSCâ brought me some real new friends for life!First of all, the fact that quickly after its inception FAIR became a hype termâ , which was probably partly even accelerated by the prominent role it played in early EOSC discussions with EC's Director General, also has its downsides. Like for the term âAIâ, everyone co-opts the term and some start watering the concept down to a bloodless caricature from what it originally meant. In the case of FAIR this includes removing the central notion of machine actionability, mis-characterising it as a standard, conflating it with âopenâ, only linking it to data sensu stricto, ignoring software, algorithms and more. In general terms, people that sometimes seem to have never read the original article [3], the most flagrant abuse of the term I have heard (obviously not from an active researcher) is this: âIf data are Findable, Accessible and Interoperable it is âautomatically' Reusable.â This is of course âswearing in FAIR churchâ as the R (principles R1â3) [3] clearly state that rich provenance and reuse conditions are critical and in particular the provenance. The decision whether (even high quality) data are fit for purpose (reuse in a particular study) is a critical step and is imho (in my humble opinion) at the basis of the reproducibility problem we currently face. Therefore, I would like to re-emphaisize here my current one liner to summarise the aim of the FAIR guiding principles: âThe Machine Knows what I meanâ. Those who feel that FAIR is too ambitious and for instance promote that âachieving F and A is enough for nowâ in my humble opinion fail to see the disruptive character of the solutions we need to make EOSC and its sister around the globe a real paradigm shift towards Open Science (OS). Or they are just trying to preserve the status quo and move incrementally at a pace they can follow.This nicely bridges to the first observation on EOSC as such. I indeed think that the first âCommunicationâ that needed 126 iterations mentioned by Jean-Claude, which happened in the same time frame as our âHLEG-1â period, was symptomatic for a basic flaw in the discussions, which haunts us still today. Conflating the âICTâ/HPC (or basic e-infrastructure) with the data and end user applications for analytics, has caused an enormous hurdle. In the entire journey of the HLEG we had to carefully navigate around this cliff and it is still a highly controversial issue today. This part was the âDunning Kruger effectâ [4] pur sang: The âother side is easyâ (because I am not hindered by any knowledge about it) and is âmore or less already doneâ (because I do not understand the complexity). This is not only true for the active researchers who cannot use the current e-infrastructure efficiently (and naturally that is âentirely the fault of the nerds who build things I do not understand or cannot operateâ), but also for e-infrastructure engineers who know everything about ICT and âthusâ (?) also about data (because âthat is just ones and zerosâ) as Jean- Claude also noted. I also believe however, that it is a mistake to completely separate e-infrastructure for the data and services layer, as the e-infrastructure should route (and understand at least at middleware level) what processes are needed on the data and how the FAIR services ârunâ. Nowadays (after many iterations) I use the diagram below (Figure 1) to explain that all three basic elements of the âInternet of FAIR Data and Servicesâ are needed. Each of them should be adorned with FAIR (machine actionable) metadata to seamlessly form a Web of FAIR Data and Services on top of the current, proven Internet backbones, thus forming the âInternet of FAIR Data and Servicesâ, eventually creating an âInternet for Social Machinesâ [5] where people and machines can both efficiently use all services, independently and in collaboration.This does absolutely not mean that the foundation (e-infrastructure) of the triangle is âtrivialâ or âcan be reused as isâ. Not only middleware, but also the crucial and fundamental concept of FDOs needs to be developed in close collaboration between data and computer experts and is largely domain-agnostic.The seamless combination will become the principle âpackageâ of information that machines (and also people) can understand and act upon. Major infrastructure builders should actually co-lead this, while domain scientists need to decide on which data formats and metadata schemes (i.e. FAIR Implementation Profiles [9]) should be built on this basic schema.Together with the Dunning Kruger effect, too many overlapping and redundant projects supporting the talking/meeting/landscaping, re-landscaping and re-re landscaping' has resulted in what I became to call the âEOSC is a bigger Me syndromeâ. On the one hand, countless people voluntarily invested (and still invest) their time in the development of the EOSC, but others seem to only see EOSC as âyet another way to collect EC funding for their current solutions that are in my opinion not future- and OS proof. This misbalance between people investing their own time and effort based on intrinsic motivation and vision and on the other hand the âreliance on EC subsidyâ caused a dichotomy during the scoping years of EOSC between disruptive and âpreservativeâ approaches. The heavy reliance on EC subsidy also largely ignored the subsidiarity principle [10] and the fact that 90% of the eventual infrastructures and services that we need for EOSC will be paid by the MSs. Also data and research intensive industry was largely kept out of the loop, which was another mistake I have frequently pointed out. This helped to create and sustain the âBrussels Bubbleâ that Jean-Claude described. The Association will hopefully reverse that trend.Finally, the influence on the HLEG report of the then-commissioner was rather profound. The report was not only delayed almost 6 months after its proposed publication version, but there is also a nice additional âuntold storyâ here: The originally proposed title of the report was: âA Cloud on the 2020 Horizonâ. In my original foreword I explained the slightly âgloomingâ connotation of that title. When the report was finally approved, it appeared that the title had been unilaterally changed into âRealising A European Open Science Cloud [11]â. Not only did I have to hastily change my foreword (because it made no sense anymore) but also, my notorious statement that the âresultâ should neither be âEuropeanâ (only), nor Open (only) nor (only) for Science and certainly not (just) a âCloudâ was entirely ignored in changing that title. But it again emphasises the âThis is an EC thingâ context, with the associated risk for confiscation of the concept by the âusual suspectsâ in EC subsidy land. However, I feel after three years of intensive deliberations, which may be considered lightning fast on the geological time scale, see George's reaction, we can conclude that most of the original HLEG recommendations are well-represented in the basic guiding documents of the EOSC Association, which makes me a happy man at the end of this crazy year.That leads me to the final observation: As a result of the (quote from Jean-Claude): ânon-paper seen as the political turning point in support of EOSCâ [12], GO FAIR (Global Open FAIR) [13] was started, originally by Germany and The Netherlands and soon joined by France as a temporary âkick-startâ, bottom-up approach to accelerate EOSC (see also recommendation I-2.1. in the HLEG report, annex 1).Soon, GO FAIR became really global and the agile modus operandi of practical Implementation Networks yielded a number of crucial approaches to speed up the adoption of the FAIR guiding principles and the hourglass approach [14]. Now, late 2020, when the EOSC Association is a fact, GO FAIR (1.0) has achieved its goals (early implementation steps) and we need to reflect on its future. Next to the intrinsic value of the active GO FAIR IN community [15] as such, several particular assets that I need to mention here are the development of the FAIR Implementation Profile and Metadata4Machines approach, the development of easy to install FAIR data points for open, FAIR metadata publication and indexing, and last but not least the international effort (involving many players, also outside the direct GO FAIR initiative) to develop the minimal specs of the FDO framework [7] in a more specified form than when coined in the FAIR expert group report [5]. These assets (all open source and open access) can be carried over, not only to EOSC, but will have much wider, international, impact most likely leading to a continuation of GO FAIR (2.0) beyond its original time scope, namely three years, the predicted time it would take to complete the international policy and bureaucracy process to reach the status of a formal association as we have today. I hope the leaders of the Association will optimally learn from the successes and failures and near-road-accidents of the last three years and see EOSC as the European contribution to a âGlobal Open Science Commonsâ, also known as the Internet of FAIR Data and Services, in full, open collaboration with the international organisations that are now joining forces in the Data Together initiative [16]. After all, the major challenges we face are global, so is the research needed to face them and so are the solutions we hope to fiend. I fully trust the current leadership of the association to make that vision reality.Policy recommendationsGovernance recommendationsImplementation recommendations
Explainable Artificial Intelligence (XAI) generates explanations which are used by regulators to audit the responsibility in case of any catastrophic failure. These explanations are currently stored in centralized systems. However, due to lack of security and traceability in centralized systems, the respective owner may temper the explanations for his convenience in order to avoid any penalty. Nowadays, Blockchain has emerged as one of the promising technologies that might overcome the security limitations. Hence, in this paper, we propose a novel Blockchain based framework for proof-of-authenticity pertaining to XAI decisions. The framework stores the explanations in InterPlanetary File System (IPFS) due to storage limitations of Ethereum Blockchain. Further, a Smart Contract is designed and deployed in order to supervise the storage and retrieval of explanations from Ethereum Blockchain. Furthermore, to induce cryptographic security in the network, an explanation's hash is calculated and stored in Blockchain too. Lastly, we perform the cost and security analysis of our proposed system.
Research and development of novel molecular compounds in the pharmaceutical industry can be highly costly. Lack of confidentiality can prevent a product from being patented or commercialized. As an effect, cross-organizational collaboration is virtually non-existent. In this paper, we introduce a blockchain-based solution to the collaborative drug discovery problem so that participants can maintain full ownership of the asset and upload partial information about molecules without revealing the molecule itself. A prototype is also implemented using the blockchain technology Hyperledger Fabric and analyzed from security and performance perspectives. The prototype provides a set of functionalities that makes sure that ownership is maintained, integrity is protected, and critical information remains confidential. From a performance perspective, it provides a good throughput and latency in the order of milliseconds. However, further improvements could be done to the scalability of the syst em.
Open access
Scientific Computing and Data Management
Innovative Microfluidic and Catalytic Techniques Innovation
Jan 1, 2021¡Proceedings of the ... Annual Hawaii International Conference on System Sciences/Proceedings of the Annual Hawaii International Conference on System Sciences
Raiane Coelho, Regina Braga, JosÊ David, Mårio A. R. Dantas ¡ 6 authors
Nowadays, scientific experiments are conducted collaboratively. In collaborative scientific experiments, we must consider aspects such as interoperability, privacy, and trust in shared data to allow the reproducibility of the results. A critical aspect associated with a scientific process is its provenance information, which can be defined as the origin or lineage of the data that helps understand the scientific experiment results. Another concern when conducting collaborative experiments is confidentiality, considering that only authorized personnel can share or view results. In this paper, we propose BlockFlow, a blockchain-based architecture, to bring reliability to the collaborative research, considering the capture, storage, and analysis of provenance data related to a scientific ecosystem platform (E-SECO).
Peer-review is a necessary and essential quality control step for scientific publications but lacks proper incentives. Indeed, the process, which is very costly in terms of time and intellectual investment, not only is not remunerated by the journals but is also not openly recognized by the academic community as a relevant scientific output for a researcher. Therefore, scientific dissemination is affected in timeliness, quality, and fairness. Here, to solve this issue, we propose a blockchain-based incentive system that rewards scientists for peer-reviewing other scientists' work and that builds up trust and reputation. We designed a privacy-oriented protocol of smart contracts called Ants-Review that allows authors to issue a bounty for open anonymous peer-reviews on Ethereum. If requirements are met, peer-reviews will be accepted and paid by the approver proportionally to their assessed quality. To promote ethical behavior and inclusiveness the system implements a gamified mechanism that allows the whole community to evaluate the peer-reviews and vote for the best ones.
David F. Ferraiolo, Joanna F. DeFranco, D. Richard Kuhn, Joshua Roberts
Distributed systems have always presented complex challenges, and technology trends are in many ways making the software designer's job more difficult. In particular, today's systems must successfully handle.
Open access
Scientific Computing and Data Management
Blockchain Technology Applications and Security
Innovative Microfluidic and Catalytic Techniques Innovation
Kevin Wittek, Dominik Krakau, Neslihan Wittek, James H. Lawton ¡ 5 authors
Proof of Existence as a blockchain service has first been published in 2013 as a public notary service on the Bitcoin network and can be used to verify the existence of a particular file in a specific point of time without sharing the file or its content itself. This service is also available on the Ethereum based bloxberg network, a decentralized research infrastructure that is governed, operated and developed by an international consortium of research facilities. Since it is desirable to integrate the creation of this proof tightly into the research workflow, namely the acquisition and processing of research data, we show a simple to integrate MATLAB extension based solution with the concept being applicable to other programming languages and environments as well.
Pingcheng Ruan, Tien Tuan Anh Dinh, Qian Lin, Meihui Zhang ¡ 6 authors
The success of Bitcoin and other cryptocurrencies bring enormous interest to blockchains. A blockchain system implements a tamper-evident ledger for recording transactions that modify some global states. The system captures the entire evolution history of the states. The management of that history, also known as data provenance or lineage, has been studied extensively in database systems. However, querying data history in existing blockchains can only be done by replaying all transactions. This approach is feasible for large-scale, offline analysis, but is not suitable for online transaction processing. We present LineageChain, a fine-grained, secure, and efficient provenance system for blockchains. LineageChain exposes provenance information to smart contracts via simple interfaces, thereby enabling a new class of blockchain applications whose execution logics depend on provenance information at runtime. LineageChain captures provenance during contract execution and stores it in a Merkle tree. LineageChain provides a novel skip list index that supports efficient provenance queries. We have implemented LineageChain on top of Hyperledger Fabric and a blockchainoptimized storage system called ForkBase. We conduct extensive evaluation, demonstrating benefits of LineageChain, its efficient querying, and its small storage overhead.
For many applications, data are worthy only if they are trustworthy. The concept of trust is sometimes elusive, and yet it is fundamental in data management. Even when not expressed explicitly, the correctness of computations and reliability of applications depend on trustworthy management of the data. These notions received new attention with the advent of blockchain and distributed ledger technology.
Abhishekh Patil, Amit Kumar Jha, Mohammed Moin Mulla, D. G. Narayan ¡ 5 authors
Cloud forensics investigates the crime committed over cloud infrastructures like SLA-violations and storage privacy. Cloud storage forensics is the process of recording the history of the creation and operations performed on a cloud data object and investing it. Secure data provenance in the Cloud is crucial for data accountability, forensics, and privacy. Towards this, we present a Cloud-based data provenance framework using Blockchain, which traces data record operations and generates provenance data. Initially, we design a dropbox like application using AWS S3 storage. The application creates a cloud storage application for the students and faculty of the university, thereby making the storage and sharing of work and resources efficient. Later, we design a data provenance mechanism for confidential files of users using Ethereum blockchain. We also evaluate the proposed system using performance parameters like query and transaction latency by varying the load and number of nodes of the blockchain network.
Raiane Coelho, Regina Braga, JosÊ David, Mårio A. R. Dantas ¡ 6 authors
With increasingly complex activities, scientific workflows are becoming more data-intensive. In this context, may require a collaborative, distributed or high performance (HPC) environment such as grids or clouds for their execution. Considering its extensibility feature, resources pool and pay-to-use, cloud computing environments have been increasingly adopted. Scientists are formulating their scientific experiments in a collaborative way, provisioning resources (software, hardware) and managing large volumes of data, based on cloud infrastructures. In data-driven collaborative scientific experiments, aspects such interoperability, privacy and trust in shared provenance data should be considered to allow the reproducibility of the results. In this paper, we present the BlockFlow architecture, which aims to bring trust to scientists of a scientific ecosystem platform (E-SECO) in the execution of their collaborative scientific experiments on cloud platforms.
Gracie Carter, Ben Chevellereau, Hossain Shahriar, Sweta Sneha
The healthcare system in the United States is unique. From payor to provider, patients have the freedom of choice. This creates a complicated and profitable paradigm of care. Legislation defines government expectations of data exchange; however, the methods are left to the discretion of the stakeholders. Today, devices and programs are not built to unified standards, thus they do not share data easily. This communication between software is known as interoperability. We address the health data interoperability by leveraging Fast Health Interoperable Resource (FHIR) standard, a viewer of FHIR called OpenPharma, and Blockchain technology. Our proof of concept, called "OpenPharma Blockchain on FHIR" (OBF), is interoperable by design and grants clinicians access to patient records using a combination of data standards, distributed applications, patient-driven identity management, and the Ethereum blockchain. OBF is a trustless, secure, decentralized, and vendor-independent method for information exchange. It is easy to implement and places the control of records with the patients.
Marten Sigwart, Michael Borkowski, Marco Peise, Stefan Schulte ¡ 5 authors
Abstract As data collected and provided by Internet of Things (IoT) devices power an ever-growing number of applications and services, it is crucial that this data can be trusted. Data provenance solutions combined with blockchain technology are one way to make data more trustworthy by providing tamper-proof information about the origin and history of data records. However, current blockchain-based solutions for data provenance fail to take the heterogeneous nature of IoT applications and their data into account. In this work, we identify functional and non-functional requirements for a secure and extensible IoT data provenance framework, and conceptualise the framework as a layered architecture. Evaluating the framework using a proof-of-concept implementation based on Ethereum smart contracts, we conclude that our framework can be used to realise data provenance concepts for a wide range of IoT use cases. While blockchain technology generally poses constraints on scalability and privacy, we discuss multiple solutions aiming to overcome these issues.