In this paper, we propose a novel framework for a scholarly journal, a token-curated registry (TCR). This model originates in the field of blockchain and cryptoeconomics and is essentially a decentralized system where tokens (digital currency) are used to incentivize quality curation of information. TCR is an automated way to create lists of any kind where decisions (whether to include N or not) are made through voting that brings benefit or loss to voters. In an academic journal, TCR could act as a tool to introduce community-driven decisions on papers to be published, thus encouraging more active participation of authors and reviewers in editorial policy and elaborating the idea of a journal as a club. TCR could also provide a novel solution to the problems of editorial bias and the lack of rewards/incentives for reviewers. In the paper, we discuss core principles of TCR, its technological and cultural foundations, and finally analyze the risks and challenges it could bring to scholarly publishing.
R. P. Jagadeesh Chandra Bose, Kanchanjot Kaur Phokela, Vikrant Kaulgud, Sanjay Podder
There has been a considerable shift in the way how software is built and delivered today. Most deployed software systems in modern times are created by (autonomous) distributed teams in heterogeneous environments making use of many artifacts, such as externally developed libraries, drawn from a variety of disparate sources. Stakeholders such as developers, managers, and clients across the software delivery value chain are interested in gaining insights such as how and why an artifact came to where it is, what other artifacts are related to it, and who else is using this. Software provenance encompasses the origins of artifacts, their evolution, and usage and is critical for comprehending, managing, decision-making, and analyzing software quality, processes, people, issues etc. In this paper, we propose an extensible framework based on standard provenance model specifications and blockchain technology for capturing, storing, exploring, and analyzing software provenance data. Our framework (i) enhances trustworthiness of provenance data (ii) uncovers non-trivial insights through inferences and reasoning, and (iii) enables interactive visualization of provenance insights. We demonstrate the utility of the proposed framework using open source project data.
Mélanie Clément‐Fontaine, Roberto Di Cosmo, Bastien Guerry, Patrick Moreau · 5 authors
Software is a hybrid object in the world research as it is equally a driving force (as a tool), a result (as proof of the existence of a solution) and an object of study (as an artefact). This specific status means we need to define strategies, tools and procedures which are adapted to the various issues it raises. These include the citation of contributions to software design and production, the reproducibility of research results involving software and the wider usage and long-term sustainability of the software heritage created. This opportunity note by the Committee for Open Science's Free Software and Open Source Project Group describes the issues at stake and formulates actionable recommendations.
Abstract. Data sharing and collaboration are critical to solving large scale problems. The prevailing soil data-sharing model is based on different groups sending their data to a lead party. This model is of a centralised nature and, consequently, results in the participants ceding their control and governance over their data to the lead party. Here we explore the use of a distributed ledger (blockchain) to solve the aforementioned issues. We explain what a blockchain is and some of its characteristics to then describe some features of a blockchain that makes it an interesting candidate for an inter-institutional database. Finally, we describe the potential use case of developing a global soil spectral library with multiple, independent international institutions constituting the network.
The potentiality of Blockchain technology is widespread and applied to diverse fields. Blockchain is a distributed ledger of transactions that store immutable records in chronological order in an append-only mode. Hence, humongous data is stored on the blockchain and will continuously expand over time. Blockchain has been rapidly adopted by many businesses for storing the provenance data because of its salient features like immutability, robustness and tamperproof. Blockchain stores data provenance as transactions that are collected from sources like a centralized cloud or decentralized cloud that helps in identifying cybercrimes. This paper emphasizes on the different approaches of querying the data provenance transactions stored in Ethereum Blockchain based on various search parameters using REST API web services. The approach not only queries based on the first-class data elements like blocks, transactions, account address and contract address but also queries based on the provenance data stored on the Ethereum Blockchain explained with a use case LegalProv.
With data intensive computing helping advance state-of-the-art in varied fields, data provenance and lineage continue to remain formidable challenges in assisting with integrity and reproducibility in research and applications. This is particularly challenging for distributed scenarios, where data may be originating from decentralized sources without any centralized control by a single trusted entity. To date most of the data provenance systems are specific to particular domains, and are often centralized. Distributed ledgers such as blockchains have proved quite popular and effective in addressing trust and consensus without central control. There are a few recent proposals to employ blockchains for data provenance, however, they rely on currency in order to propose transactions using public blockchains.\n\nWe present HyperProv, a general framework for data provenance based on the permissioned blockchain Hyperledger Fabric (HLF), and to the best of our knowledge, the first provenance system that is ported to ARM based devices such as Raspberry Pi (RPi). HyperProv records the operation history and data lineage by tracking checksums, editors, timestamps, data pointers, dependencies, and more. Provenance data is retrieved and stored through a NodeJS client library to simplify interactions with the blockchain. HyperProv has a set of built-in queries using smart contracts that enable lightweight retrieval of large collections of provenance data. We evaluate the throughput, latency and resource consumption of HyperProv on x86-64 desktop machines, as well as RPi, demonstrating the feasibility of using HyperProv on RPi for tamperproof data provenance, useful in particular for Internet of Things use cases.
Maribel Acosta, Tim Berners‐Lee, Stefan Dietze, Anastasia Dimou · 9 authors
Decentralised data solutions bring their own sets of capabilities, requirements and issues not necessarily present in centralised solutions. In order to compare the properties of different approaches or tools for management of decentralised data, it is important to have a common evaluation framework. We present a set of dimensions relevant to data management in decentralised contexts and use them to define principles extending the FAIR framework, initially developed for open research data. By characterising a range of different data solutions or approaches by how TRusted, Autonomous, Distributed and dEcentralised, in addition to how Findable, Accessible, Interoperable and Reusable, they are, we show that our FAIR TRADE framework is useful for describing and evaluating the management of decentralised data solutions, and aim to contribute to the development of best practice in a developing field.
The Blockchain technology was initially adopted to implement various cryptocurrencies. Currently, Blockchain is foreseen as a general purpose technology with a huge potential in many areas. Blockchain-based applications have inherent characteristics like authenticity, immutability and consensus. Beyond that, records stored on Blockchain ledger can be accessed any time and from any location. Blockchain has a great potential for managing and maintaining educational records. This paper presents a Blockchain-based Educational Record Repository (BcER2) that manages and distributes educational assets for academic and industry professionals. The BcER2 system allows educational records like e-diplomas and e-certificates to be securely and seamless transferred, shared and distributed by parties.
Blockchain, or distributed ledger technology, has been an increasingly common topic in technical circles over the past several years. You may have read one of the thousands of articles documenting ...
Key points Digital Science's paper is one of the first looking at the application of blockchain technology in scholarly publishing. Wholesale use of blockchain technologies is suggested as a possible replacement for scholarly publishers. There remain questions around the adoption of blockchain technologies, including privacy, researcher support, and fraudulent use. Blockchain technologies may provide a new means of understanding problems and customers' evolving expectations, but careful consideration is required of whether blockchain is the best solution.
Krzysztof Janowicz, Blake Regalia, Pascal Hitzler, Gengchen Mai · 8 authors
Distributed ledger technologies such as blockchains and smart contracts have the potential to transform many sectors ranging from the handling of health records to real estate. Here we discuss the value proposition of these technologies and cryptocurrencies for science in general and academic publishing in specific. We outline concrete use cases, provide an informal model of how the Semantic Web journal's peer-review workflow could benefit from distributed ledger technologies, and also point out challenges in implementing such a setup.
Purpose The purpose of this paper is to employ the case of Organization for Economic Cooperation and Development (OECD) data repositories to examine the potential of blockchain technology in the context of addressing basic contemporary societal concerns, such as transparency, accountability and trust in the policymaking process. Current approaches to sharing data employ standardized metadata, in which the provider of the service is assumed to be a trusted party. However, derived data, analytic processes or links from policies, are in many cases not shared in the same form, thus breaking the provenance trace and making the repetition of analysis conducted in the past difficult. Similarly, it becomes tricky to test whether certain conditions justifying policies implemented still apply. A higher level of reuse would require a decentralized approach to sharing both data and analytic scripts and software. This could be supported by a combination of blockchain and decentralized file system technology. Design/methodology/approach The findings presented in this paper have been derived from an analysis of a case study, i.e., analytics using data made available by the OECD. The set of data the OECD provides is vast and is used broadly. The argument is structured as follows. First, current issues and topics shaping the debate on blockchain are outlined. Then, a redefinition of the main artifacts on which some simple or convoluted analytic results are based is revised for some concrete purposes. The requirements on provenance, trust and repeatability are discussed with regards to the architecture proposed, and a proof of concept using smart contracts is used for reasoning on relevant scenarios. Findings A combination of decentralized file systems and an open blockchain such as Ethereum supporting smart contracts can ascertain that the set of artifacts used for the analytics is shared. This enables the sequence underlying the successive stages of research and/or policymaking to be preserved. This suggests that, in turn, and ex post , it becomes possible to test whether evidence supporting certain findings and/or policy decisions still hold. Moreover, unlike traditional databases, blockchain technology makes it possible that immutable records can be stored. This means that the artifacts can be used for further exploitation or repetition of results. In practical terms, the use of blockchain technology creates the opportunity to enhance the evidence-based approach to policy design and policy recommendations that the OECD fosters. That is, it might enable the stakeholders not only to use the data available in the OECD repositories but also to assess corrections to a given policy strategy or modify its scope. Research limitations/implications Blockchains and related technologies are still maturing, and several questions related to their use and potential remain underexplored. Several issues require particular consideration in future research, including anonymity, scalability and stability of the data repository. This research took as example OECD data repositories, precisely to make the point that more research and more dialogue between the research and policymaking community is needed to embrace the challenges and opportunities blockchain technology generates. Several questions that this research prompts have not been addressed. For instance, the question of how the sharing economy concept for the specifics of the case could be employed in the context of blockchain has not been dealt with. Practical implications The practical implications of the research presented here can be summarized in two ways. On the one hand, by suggesting how a combination of decentralized file systems and an open blockchain, such as Ethereum supporting smart contracts, can ascertain that artifacts are shared, this paper paves the way toward a discussion on how to make this approach and solution reality. The approach and architecture proposed in this paper would provide a way to increase the scope of the reuse of statistical data and results and thus would improve the effectiveness of decision making as well as the transparency of the evidence supporting policy. Social implications Decentralizing analytic artifacts will add to existing open data practices an additional layer of benefits for different actors, including but not limited to policymakers, journalists, analysts and/or researchers without the need to establish centrally managed institutions. Moreover, due to the degree of decentralization and absence of a single-entry point, the vulnerability of data repositories to cyberthreats might be reduced. Simultaneously, by ensuring that artifacts derived from data based in those distributed depositories are made immutable therein, full reproducibility of conclusions concerning the data is possible. In the field of data-driven policymaking processes, it might allow policymakers to devise more accurate ways of addressing pressing issues and challenges. Originality/value This paper offers the first blueprint of a form of sharing that complements open data practices with the decentralized approach of blockchain and decentralized file systems. The case of OECD data repositories is used to highlight that while data storing is important, the real added value of blockchain technology rests in the possible change on how we use the data and data sets in the repositories. It would eventually enable a more transparent and actionable approach to linking policy up with the supporting evidence. From a different angle, throughout the paper the case is made that rather than simply data, artifacts from conducted analyses should be made persistent in a blockchain. What is at stake is the full reproducibility of conclusions based on a given set of data, coupled with the possibility of ex post testing the validity of the assumptions and evidence underlying those conclusions.
This paper offers an overview of the highlights of the NFAIS Conference, Blockchain for Scholarly Publishing, that was held in Alexandria, VA from May 15–16, 2018. The goal of the conference was to take a close look at the initiatives that have emerged as a result of the increasing global acceptance of blockchain technology. This technology, chiefly known as the foundation of Bitcoin and originally introduced as a means of securely managing cryptocurrency, has proven to have practical applications beyond finance. The basic technology is that of a distributed ledger and it is being broadly-adopted by multiple industries, including the scholarly publishing community. The capabilities of this new technology are prompting a direct exchange among stakeholders, as blockchain promises a more structured, decentralized, and immutably secure approach that has the potential to significantly impact researcher workflows - from data collection to peer review to access and published work. The technology inspires passion - there are those who believe that it will ultimately transform our lives while others are completely skeptical. The NFAIS conference provided a look at both sides of the coin (no pun intended).
This article presents a new method for managing digital reuse rights of research data, which leverages technologies such as the blockchain and smart contracts. This allows, on one hand, the creation of a permanent record on the agreements between the authors of the data and the reusers, with the possibility of verifying compliance at any time, and on the other hand, a higher level of granularity on defining the conditions of reuse. A practical implementation of such a workflow using the Solidity smart contract language is included, along with a brief analysis over the Ethereum blockchain network.
The purpose of this thesis was to investigate and study the various issues faced by educational and technological researchers while raising the funds for their respective projects and the issues faced by the fund’s providers. Multiple existing traditional fundraising platforms were identified, and their advantages and disadvantages were studied to check if it was suitable for educational and technological researchers to carry on their funding campaign using the existing platforms. Finally, the goal was to develop a decentralized research funding application which would replace the existing traditional methods of raising funds by providing the researchers the ability to create a fundraising campaign on Ethereum blockchain while ensuring the transparent and auditable usage of the funds provided for the development of the project by the stakeholders.\n\nThe research funding application was developed and deployed to Ethereum blockchain. During the development process, the technologies used were Solidity, HTML, CSS, Javascript and React. The requirements for the Minimum Viable Product of the research funding application were finalized and the project was implemented by following the Waterfall software development model. \n\nAs a result, the requirements set for the research funding application were accomplished and the application was deployed to the blockchain and can be accessed by the general public. Furthermore, additional features such as the ability to create and manage multiple funding campaigns by a single entity were also developed successfully.
Jonathan Bell, Thomas D. LaToza, Foteini Baldmitsi, Angelos Stavrou
The scientific community is facing a crisis of reproducibility: confidence in scientific results is damaged by concerns regarding the integrity of experimental data and the analyses applied to that data. Experimental integrity can be compromised inadvertently when researchers overlook some important component of their experimental procedure, or intentionally by researchers or malicious third-parties who are biased towards ensuring a specific outcome of an experiment. The scientific community has pushed for "open science" to add transparency to the experimental process, asking researchers to publicly register their data sets and experimental procedures. We argue that the software engineering community can leverage its expertise in tracking traceability and provenance of source code and its related artifacts to simplify data management for scientists. Moreover, by leveraging smart contract and blockchain technologies, we believe that it is possible for such a system to guarantee end-to-end integrity of scientific data and results while supporting collaborative research.
Christopher Munro, Philip Couch, Jon Johnson, John Ainsworth · 5 authors
Discovery of useful relationships between scholarly assets on the web is challenging, both in terms generating the right metadata around the assets, and in connecting all relevant digital entities in chain of provenance accessible to the whole community. This paper reports the development of a framework and tools enabling scholarly asset relationships to be expressed in a standard and open way, illustrated with use-cases of discovering new knowledge across cohort studies. The framework uses Research Objects for aggregation, distributed databases for storage, and distributed ledgers for provenance. Our proposal avoids management by a single central platform or organization, instead leveraging the use of existing resources and platforms across natural partnerships. Our proposed infrastructure will support a wide range of users from system administrators to researchers.
The institutional repositories of most university libraries in Korea are not activated due to the lack of finance, operation ability, technology, etc., and this phenomenon is especially noticeable in small–medium sized university libraries. Thus, as in other countries, a shared repository that multiple libraries share, based on a library network, is required. This study categorizes the typology of shared repositories, after analyzing the operational methods of managing roles, expenses and system sharing among participants in shared repositories in Japan, the UK and the USA. Based on the findings, two models that could be applied in Korea are suggested. The first is a centralized-operation model that involves a shared-system infrastructure in which the host institution takes full charge of the digitization and registration of contents, etc. This model has the potential to be developed into a regional archive center. The second model is a decentralized-operation model that also features a shared-system infrastructure but the participating institutions customize the system individually. Here, the individual institutions perform digitization of all content, daily registration, etc., which can be applied as a test bed for the creation of new technology, considering the potential for independent operation in the future.
PIMMS (Portable Infrastructure for the Metafor Metadata System) provides institutions with tools to capture information about the workflow of running simulations from the design of experiments to the implementation of experiments via simulations running models. PIMMS uses the Metafor methodology for simulation documentation which consists of a common information model (CIM), a set of controlled vocabularies (CV) and software tools. PIMMS software tools provide for the creation and consumption of CIM content via a web infrastructure and portal.PIMMS will refactor the "CMIP5 questionnaire" metadata management tool, that is collecting climate model metadata for the CMIP5 model inter-comparison project, so that it can be more easily portable into stand alone installations within the university environment and customised to address the specific requirements of individual research groups. Initial model descriptions may take time to complete but once they have been cre ated the PIMMS infrastructure can be used to document subsequent variations by describing only those elements that are changed. An established PIMMS infrastructure will fit seamlessly into the research metadata workflow and significantly reduce subsequent documentation effort. The key to the customisation of PIMMS is in the modularity of its tools and the clear separation of structure (CIM) from content (CV). The PIMMS project will extend the CMIP5 controlled vocabulary to encompass descriptions of paleoclimate models and will also demonstrate how the CIM can be used to document an Integrated Assessment Model (IAM). This proof of concept prototype will create a new controlled vocabulary in collaboration with Ermitage and use it to reconfigure PIMMS to collect metadata in a different discipline. PIMMS will further explore how the CV that is used to configure PIMMS may be of further use to our stake holders and the wider JISC community through the development of the Uni versity of Cambridge chemicaltagger tool. PIMMS will provide a local portal so that research groups can view and search their own content, as well as publish their metadata content to institutional, national and international services. In addition PIMMS will also include data node software so that data documented with PIMMS can also be published to the web, both locally, and to national and international services.
With the rapid growth in the number of electronic resources available via the Internet, a variety of methods have been developed to organize and access these objects. Librarians, scholars and computing engineers each have applied their own techniques to the process. Catalogers envision a "super-catalog"; scholars spawn encoded texts; computer engineers design robot-generated indexes. Each method has its strengths and its weaknesses, and each group casts a wary eye upon competing systems. In fact, it is unrealistic to believe that the organizational system of any one group will be adopted by all players in the electronic arena. While general principles of organization should apply to all approaches, cataloging the Internet is inherently different from cataloging a library. Libraries are systematically developing collections of primarily fixed objects usually under the control of one institution or agency. The Internet is more analogous to the ubiquitous ebb and flow of information and services appearing throughout society. The structure and content are not systematically developed and stable; there is no single user group or purpose for the Internet; and there is no single controlling agency. Thus, there is no one community, be it scholars, computing engineers or librarians, that is clearly vested with authority to decide upon the best means to bring control over this electronic chaos. Considering that the Internet has always been a decentralized initiative, it is not surprising that the efforts to organize it have been similarly independent and uncoordinated. Analogous to a loosely coupled communication system, each community working to organize the Internet is somewhat responsive but essentially autonomous. It is likely that each of these autonomous groups will continue to develop methods for description and access that best meet their own perceived needs. At best, cooperative efforts may result in some common understanding of a core set of descriptive data; but it is unlikely that the application, structure and use of that data will be identical within all communities. It is important, therefore, that each group recognize the contributions of the others and that together they provide bibliographic control methods that can be layered, interchanged and translated within a broad but loosely coupled system of organization. In this way, each community can continue to develop methods compatible with its own users' needs, while taking advantage of data and systems created by other groups. Remote electronic resources have engendered discussion of the catalog's scope and functions, the concept of a collection and the appropriateness of including bibliographic surrogates for remote resources in a local catalog. Traditionally, library catalogs have provided macro level access to whole items in the collection, relying on other bibliographic tools, such as indexes and bibliographies, to provide micro level access to bodies of literature and parts of items. One criticism of this distributed bibliographic system is its failure to provide access to everything in the library collection through one totally integrated catalog. More sophisticated computer technology offers new possibilities for correcting this problem. Electronic indexes and databases mounted on the local computer provide, if not a truly integrated catalog, a reasonably integrated interface that allows access to several bibliographic tools from one terminal. Work continues to develop more sophisticated layering systems that offer searchers the convenience of accessing simultaneously the OPAC and several other databases with a single search query. While providing the illusion of searching a single database, these systems could use both front-end interfaces that convert differing search commands into one common command language and multi-language thesauri that translate terms used by one database into those used in another. These techniques of layering, exchanging and translating data allow us to avoid the time, expense and proprietary problems involved with actually creating one large integrated catalog. This is another example of a loosely coupled organizational system and provides a model for the more complex goal of organizing the Internet. Due to the economic constraints of the past decade, most libraries have shifted their collection development focus from exclusive "ownership" to providing "access," a shift that has not only blurred the concept of "collection," but has caused many librarians and information specialists to lose sight of the beneficial service provided by the process of collection development. It is essential and desirable that the confining parameters that define a collection be expanded to accommodate documents that are not owned and physically housed within the library's walls; however, it is just as essential and desirable to continue the value-added concept of a collection as interrelated parts selected with a purpose. Remotely accessed electronic resources should not be excluded from the collection development scrutiny accorded documents in all other formats. In light of discussions on scholarly communication and the problems of electronic "publishing" without peer review, the inclusion of Internet resources in the library catalog based upon appropriate selection criteria may be even more necessary than in the past to assist users in finding relevant resources. The current organization of electronic resources can be described at two levels: the local agency's catalog and catalogs of Internet resources beyond the auspices of any one library. At level one, a description of the resource is contained in the local library catalog, along with bibliographic surrogates for all other materials for which that library provides access, artifactual and electronic. The Anglo-American Cataloging Rules have proved adaptable to new formats and have been modified, interpreted and supplemented to accommodate remote electronic resources. The existing MARC record also has been revised to provide special fields and coded data elements for the description and access of remote electronic resources. Specifically, field 856 has been added to convey information necessary to locate and access the electronic object. Thus, the current methods of cataloging continue to be used to describe resources that are in some ways fundamentally different from those currently owned by libraries. At present, however, only a portion of library catalogs are able to provide active links through the OPAC to access the electronic documents directly; but as libraries migrate to Web-based catalogs, these "hotlinks" will become more common. It is not surprising that the vanguard attempts to catalog Internet resources use current and familiar cataloging methods. But this AACR2/MARC-based cataloging system is not without problems. The MARC record is designed primarily for single object description and linear access. It cannot easily be adapted to describe adequately multi-level hypertext objects. In addition, these bibliographic records are designed to carry a large amount of carefully selected data and to allow access to much of that data. They are, therefore, detailed and complex structures which require much time to create and encode correctly. Recent efforts to save time and money by simplifying the cataloging process argue against traditional cataloging methods that create full MARC records for each electronic resource. As library administrators attempt to save money by purchasing cataloging data from outside sources, new ways of creating bibliographic records are being sought. While the library cataloging community was attempting to determine the feasibility of using traditional cataloging standards for Internet resources, scholars in the area of humanities computing concentrated on the electronic document itself and on a project to develop a text encoding scheme for complex electronic textual objects. The project became known as the Text Encoding Initiative (TEI). The TEI guidelines give recommendations on what features of the text to encode and how to encode them. They also include a chapter on creating a header imbedded within the electronic document that precedes the electronic text. This TEI header consists primarily of four parts: file description, encoding description, text profile and text revision history. Designed for multipurpose use, the TEI header provides metadata needed by librarians who will catalog the text, scholars who will use the text, and software programs that will operate on the text. The file description portion of the header is intended to serve as the electronic equivalent of the title page of a printed work and is the only required part of the header. When data elements for all four components of the header are encoded, TEI header structures may be as lengthy and complex as MARC records. Several groups within the library community are now examining the possibility of using TEI header information as a source of data in the automated creation of MARC records for electronic documents. Computer programs have been produced that convert TEI-tagged bibliographic headers into MARC formatted records. As research develops in the area of metadata formats, the process of TEI header to MARC data conversion may become the copy cataloging of tomorrow. At level two, the goal is to organize Internet resources independently of any library agency. Several means of organization and access currently exist, including separate catalogs of selected Internet resources, subject browsing lists and robot-generated search tools. The common feature of all these "second level" access tools is the exclusive focus on Internet resources. OCLC's InterCat Catalog is one example of a second level bibliographic tool designed through a cooperative library endeavor. Consisting of MARC records created by individual catalogers from many libraries, the database, with its Web interface, represents the beginnings of an Internet Union Catalog (IUC) that is not tied to any one library as information access provider. The database still relies, however, on individual catalogers creating and entering MARC records. The IUC, while one level above the local library OPAC, still contains many of the local OPAC's useful characteristics. The documents included have been selected by librarians for their quality and appropriateness; and bibliographic surrogates in the database contain the kinds of data and authority controlled access points that have proven useful in information retrieval. When the level two organizational goal extends beyond the library selection process to the entire corpus of information on the Internet, the task of providing description and access for these electronic resources and services becomes monumental. The amount of information requiring organization is beyond the existing methods and systems of individual professional catalogers, indexers and abstractors to manage either efficiently or effectively. Several communities have wrestled with the problem and developed different ways to provide access to these resources. Over the last few years, a dramatic increase in the amount of information on the Internet heightened the need for improved access methods. Savvy computer engineers realized they had to develop retrieval systems that could be mastered easily by the non-scientific community that now used the Internet so heavily. Their efforts resulted in three primary access methods: direct address, directory browsing lists and robot-generated searchable indexes. Direct address, i.e., entering the specific pathname, directory and filename (e.g., the URL), is a known object search and requires no organizational tools to locate the site. The two remaining methods, directory browsing lists and robot-generated searchable indexes, are organizational tools that are used heavily across all Internet user communities. Directory browsing lists help people find Internet resources by arranging those resources by subject. The most commonly used lists, such as Yahoo!, use an alphabetico-classed arrangement with verbal subject topics hierarchically ordered. This arrangement is especially useful for browsing, where the user can move from the general subject to the more specific topic and then connect directly to the resources assigned to that topic. Although easy to use, this subject arrangement presents several problems: there are few cross references; many users may find difficulty identifying the hierarchy under which their narrower topic might be found; and these lists soon become lengthy and unwieldy. Other browsing tools, created outside but influenced by the library community, such as Patrick's Subject Catalog and Mundie's CyberDewey, are arranged by classification number. This affords a more logical hierarchical subject approach that is not dependent on the vagaries of the alphabet for its arrangement. While remaining less tied to semantic content than the verbally based subject systems, some type of verbal clarification is required to interpret the notation correctly. The Dewey Decimal Classification appears to be the system of choice for most of these classification-based browsing tools and has the advantage of being familiar to many users, both within the United States and internationally. As with the other subject lists, however, classification lists also need a system of cross-referencing, and the subject lists can become lengthy and unwieldy. Directory browsing lists also have other limitations. Most require a great deal of human effort to collect, arrange, encode and annotate the list of resources. When the vast number of Internet resources to be organized is considered, this could be viewed as a serious handicap. The second type of search tool, the robot-generated searchable index, attempts to overcome many of the limitations of the browsing lists. Robot-generated indexes, such as Lycos and Alta Vista, rely on computer power to collect and index electronic resources and provide an interactive interface that allows the user to enter searches for specific terms or phrases. Many allow the user to perform very sophisticated searches employing Boolean operators, adjacency and proximity operators and imbedded truncation. As a result of the automated collecting and indexing process, these search engines provide access to a much larger volume of electronic resources than the browsing lists. These computer generated search tools also have their own weaknesses. One problem is that they frequently produce extensive lists of search results, requiring users to spend time examining a great number of irrelevant and often incomprehensible citations in order to find a few pertinent hits. Searchers familiar with library catalogs and bibliographic databases may well wonder at the bewildering content of many retrieved citations. Many confusing citations are caused by extracting data directly from the document without human review and by the lack of standardized descriptions and authority control for names and subjects. Thus, while directory browsing lists and robot-generated indexes have the advantage of providing immediate access to a vast array of Internet resources and are often easy to use, they lack many features, such as selectivity, descriptive citation data and authority control, that have made the more traditional bibliographic tools most useful. If robot-generated indexes provide too little descriptive information for effective and efficient document retrieval and manually created MARC and TEI headers contain too much information for rapid and inexpensive document description, is there some method of describing Internet resources that mediates between these two extremes? With this in mind, the Dublin Core data set was developed to perform this function by defining a core data set that could describe a wide range of electronic objects. If such a core data set were adopted as an established standard and incorporated into the electronic object at the time of creation, it would increase the amount and quality of available resource descriptions with a minimal amount of human intervention. These descriptions would improve the quality of index citations created by automated tools, while enabling MARC or TEI header based records to be developed with much less expense. The Dublin Core development workshop focused on data elements necessary for the discovery of document-like objects, with consideration given to mechanisms for extending the elements to meet specialized needs. The syntax of the Dublin Core was left deliberately unspecified so that the data elements could be mapped into a variety of more complex structures from the MARC format and TEI headers to some as yet undefined structure. It is this aspect of sharing and translating descriptive data that holds the most promise, for by creating a minimum standard for descriptive data that can be used as desired by each community, we have the basis for an umbrella organizational structure based on a loosely coupled system. Then each stakeholder can continue to develop and improve its own system and structures, while overarching structures can be created to exchange, translate and layer common data. While level two provides access to a much wider range of Internet resources than any level one catalog, most information about nonelectronic resources that is provided at level one is lost with level two access tools. Although it is possible to locate and search library OPACs on the Internet via telnet or a Web interface, OPAC searches must be performed as a separate activity. A single query will not access information found at-large on the Internet and in the library OPAC, because library OPAC databases are not included in searches performed by current Internet search engines. As we move forward with efforts to create minimum standards for descriptive data, develop ways for creators of electronic resources to provide descriptive information at the time of document creation and refine search engines to enable simultaneous searches of multiple databases, we should have as our goal the creation of a third level access tool — the metacatalog. We must develop search tools whereby a user can identify specific library catalogs to include in a search query of other Internet databases, much as some existing search engines allow a user to select Web, gopher, ftp, newsgroups or commercial databases for searching. The metacatalog should be able to translate and interpret MARC records, TEI headers, SGML, HTML and any future coding. There should be language translators built in, so that searchers can use search commands and view data labels in their language of choice and translate terms in subject browsers. Multiple thesauri could be employed to assist with vocabulary control across communities, and a high-level authority control component could transparently identify and link variant forms of names. In other words, the next generation metacatalogs should be able to access all relevant information seamlessly, no matter what format or language. In order to accomplish this each stakeholder community must stop trying to recast the tools created by others into ones of their own structure, but rather concentrate on developing ways to layer, exchange and translate data within a loosely-coupled organizational system.