<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<!--<?xml-stylesheet type="text/xsl" href="article.xsl"?>-->
<article article-type="review-article" dtd-version="1.1" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id journal-id-type="issn">1683-1470</journal-id>
<journal-title-group>
<journal-title>Data Science Journal</journal-title>
</journal-title-group>
<issn pub-type="epub">1683-1470</issn>
<publisher>
<publisher-name>Ubiquity Press</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5334/dsj-2019-052</article-id>
<article-categories>
<subj-group>
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>The History and Future of Data Citation in Practice</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-7723-0950</contrib-id>
<name>
<surname>Parsons</surname>
<given-names>Mark A.</given-names>
</name>
<email>parsom3@rpi.edu</email>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0003-4808-4736</contrib-id>
<name>
<surname>Duerr</surname>
<given-names>Ruth E.</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0003-0077-4738</contrib-id>
<name>
<surname>Jones</surname>
<given-names>Matthew B.</given-names>
</name>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
</contrib-group>
<aff id="aff-1"><label>1</label>Rensselaer Polytechnic Institute (RPI), US</aff>
<aff id="aff-2"><label>2</label>Ronin Institute, US</aff>
<aff id="aff-3"><label>3</label>University of California Santa Barbara, US</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2019-11-01">
<day>01</day>
<month>11</month>
<year>2019</year>
</pub-date>
<pub-date pub-type="collection">
<year>2019</year>
</pub-date>
<volume>18</volume>
<elocation-id>52</elocation-id>
<history>
<date date-type="received" iso-8601-date="2019-07-30">
<day>30</day>
<month>07</month>
<year>2019</year>
</date>
<date date-type="accepted" iso-8601-date="2019-10-04">
<day>04</day>
<month>10</month>
<year>2019</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright: &#x00A9; 2019 The Author(s)</copyright-statement>
<copyright-year>2019</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See <uri xlink:href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</uri>.</license-p>
</license>
</permissions>
<self-uri xlink:href="http://datascience.codata.org/articles/10.5334/dsj-2019-052/"/>
<abstract>
<p>In this review, we adopt the definition that &#8216;Data citation is a reference to data for the purpose of credit attribution and facilitation of access to the data&#8217; (<xref ref-type="bibr" rid="B67">TGDCSP 2013: CIDCR6</xref>). Furthermore, access should be enabled for both humans and machines (<xref ref-type="bibr" rid="B25">DCSG 2014</xref>). We use this to discuss how data citation has evolved over the last couple of decades and to highlight issues that need more research and attention.</p>
<p>Data citation is not a new concept, but it has changed and evolved considerably since the beginning of the digital age. Basic practice is now established and slowly but increasingly being implemented. Nonetheless, critical issues remain. These issues are primarily because we try to address multiple human and computational concerns with a system originally designed in a non-digital world for more limited use cases. The community is beginning to challenge past assumptions, separate the multiple concerns (credit, access, reference, provenance, impact, etc.), and apply different approaches for different use cases.</p>
</abstract>
<kwd-group>
<kwd>data citation</kwd>
<kwd>FAIR</kwd>
<kwd>credit</kwd>
<kwd>access</kwd>
<kwd>impact</kwd>
<kwd>micro-citation</kwd>
<kwd>persistent identifiers</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec>
<title>I. Introduction</title>
<p>Data citation helps make data sharing both more FAIR &#8212; findable, accessible, interoperable, and reusable (<xref ref-type="bibr" rid="B68">Wilkinson et al. 2016</xref>) &#8212; and fair. Citation helps make data more findable and accessible through current scholarly communication systems. It can aid interoperability through precise reference to data and associated services and aid reusability by providing some context of how data have been created and used. Citation also helps credit the intellectual effort necessary to create a good data set and provides accountability for that data. It recognizes important scientific contributions beyond the written publication and, therefore, makes things fairer for everyone involved in doing good science.</p>
<p>Data citation is not a new concept. It rests on a fundamental principle of the scientific method that demands recognized, verifiable, and credited evidence behind an assertion. Traditionally, this was done within the literature through citation of materials often held in special library collections, such as field or lab logs, printed books, monographs, and maps (<xref ref-type="bibr" rid="B28">Downs et al. 2015</xref>). If the data were small enough, they were simply included directly in the publication. Some fields, such as astronomy, had whole journals devoted to publishing data. The digital age and the corresponding growth in the volume and complexity of data changed all that.</p>
<p>In this review, we adopt the definition that &#8216;Data citation is a reference to data for the purpose of credit attribution and facilitation of access to the data (<xref ref-type="bibr" rid="B67">TGDCSP 2013: CIDCR6</xref>). Furthermore, access should be enabled for both humans and machines (<xref ref-type="bibr" rid="B25">DCSG 2014</xref>). We use this to discuss how data citation has evolved over the last couple of decades and to highlight issues that need more research and attention.</p>
<p>Early work illustrates the desire to build from existing systems and culture (e.g., <xref ref-type="bibr" rid="B4">Altman &amp; King 2007</xref>). Silvello (<xref ref-type="bibr" rid="B62">2018</xref>) provides an excellent, extensive review of the motivations, principles, and high-level practices of data citation. Borgman (<xref ref-type="bibr" rid="B14">2016</xref>) provides a similar review and illustrates the disconnect between bibliographic citation principles and data citation.</p>
<p>We build on this work by focusing on how the two concerns, credit and access, and the two audiences, humans and machines, can create tensions in how data citation is conceived and implemented. We review relevant literature and current activities while also drawing from our own experience working as data professionals managing data, systems, and communities for decades. We come from the perspective of observational science, where data record phenomena as they occur and cannot be repeated. This makes precise citation more necessary and more challenging. Our experience is primarily in Earth, environmental, and space science, but we believe our recommendations apply broadly.</p>
</sec>
<sec>
<title>II. History</title>
<p>The modern concept of data citation emerged in the late 1990s. For example, one of the authors was involved in an effort at that time, where NASA&#8217;s Earth science archives (the DAACs) agreed to a common approach to data citation. The US Geological Survey also proposed guidelines (<xref ref-type="bibr" rid="B10">Berquist Jr. 1999</xref>). But neither of these approaches were broadly adopted even within NASA and USGS. Part of the issue was that journals were still developing standards on how to cite electronic resources in general. In the analog era, the intangible concepts of an article were manifest in a physical object, but as we moved into the digital age, we needed &#8220;to think not of one space (the physical, paper space) but of three &#8216;spaces&#8217; in which &#8216;the same&#8217; articles appear:</p>
<list list-type="bullet">
<list-item><p>Information space = the work as intangible entity (ideas)</p></list-item>
<list-item><p>Cyberspace = digital manifestation (electronic, made of bits)</p></list-item>
<list-item><p>&#8216;Paper space&#8217; = physical manifestation (cellulose and ink, made of atoms)&#8221; (<xref ref-type="bibr" rid="B58">Paskin 2000: 2</xref>).</p></list-item>
</list>
<p>This added more complexity to the concept of citation, and the problem was exacerbated by the impermanence of web references (e.g., <xref ref-type="bibr" rid="B49">Lawrence et al. 2001</xref>). The Digital Object Identifier (DOI) first emerged in the year 2000,<xref ref-type="fn" rid="n1">1</xref> even though the underlying Handle system was almost as old as the web (<xref ref-type="bibr" rid="B39">Kahn &amp; Wilensky 1995</xref>). Overall, in those early digital years, data were rarely cited, and if they were, the mechanisms were erratic and inconsistent.</p>
<p>From the mid-2000s, there was a growing consensus on the use of registered, resolvable, authority-based persistent identifiers (PIDs) &#8212; not only for papers but also other digital artifacts, notably data. DOIs emerged as the PID of choice for many publishers and data repositories, but there are other popular choices, such as Archival Resource Keys (ARKs)<xref ref-type="fn" rid="n2">2</xref> or compact URIs (CURIEs) resolved through metaresolvers like Name to Thing<xref ref-type="fn" rid="n3">3</xref> or <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://identifiers.org">identifiers.org</ext-link>. Indeed, local identifiers (accession numbers) have been used for centuries for internal management, and even external referencing, especially for biocollections. Moreover, digital entities (e.g., computer files), physical entities (e.g., rock samples), living things (e.g., wildlife), and descriptive entities (e.g., mitosis) have different requirements for identifiers (<xref ref-type="bibr" rid="B33">Guralnick et al. 2015</xref>). McMurry et al. (<xref ref-type="bibr" rid="B52">2017</xref>) provide a good contemporary review of identifiers and how to use them, and the RDA Data Fabric Interest Group has developed a set of assertions about the nature, creation and implementation of PIDs (<xref ref-type="bibr" rid="B69">Wittenburg et al. 2017</xref>).</p>
<p>At the same time, there was an emerging argument that data should be &#8216;published&#8217; in a manner akin to scientific literature and cited accordingly (<xref ref-type="bibr" rid="B18">Callaghan et al. 2009</xref>; <xref ref-type="bibr" rid="B21">Costello 2009</xref>; <xref ref-type="bibr" rid="B43">Klump et al. 2006</xref>; <xref ref-type="bibr" rid="B48">Lawrence et al. 2011</xref>). Data started getting DOIs in 2004 through a pilot project in Germany, and DataCite was established in December 2009, with the explicit global mission of minting DOIs for data (<xref ref-type="bibr" rid="B44">Klump et al. 2015</xref>). Even as the &#8216;publication&#8217; paradigm for data was questioned (<xref ref-type="bibr" rid="B56">Parsons &amp; Fox 2013</xref>; <xref ref-type="bibr" rid="B61">Schopf 2012</xref>), citation was broadly supported by the library and data management community. This support culminated with the broadly endorsed <italic>Joint Declaration of Data Citation Principles</italic> in 2014, which defined the core purposes and general practices of data citation (<xref ref-type="bibr" rid="B25">DCSG 2014</xref>). But this was an agreement of the library, data, and information science communities. The research community remained largely unaware, and studies have revealed that data citation remains an infrequent and inconsistent process (<xref ref-type="bibr" rid="B37">Howison &amp; Bullard 2016</xref>; <xref ref-type="bibr" rid="B53">Mooney &amp; Newton 2012</xref>; <xref ref-type="bibr" rid="B51">Mayernik et al. 2016</xref>; <xref ref-type="bibr" rid="B62">Silvello 2018</xref>).</p>
<p>Part of the reason for infrequent and inconsistent data citation is that, from the researcher&#8217;s point of view, making data FAIR implies sharing and reuse and therefore effort by the researcher. Based on our collective experience managing multiple data archives and networks, large scale community data, like satellite imagery, may be reused a lot, but much data from more-localized, research collections may never be shared or reused until the broad community recognizes that aggregating these data globally is necessary to make further progress (<xref ref-type="bibr" rid="B55">Parsons et al. 2008</xref>; <xref ref-type="bibr" rid="B5">Baker &amp; Yarmey 2009</xref>). This takes time and effort. For example, starting in 1995 with the creation of the National Center for Ecological Analysis and Synthesis (NCEAS) (<xref ref-type="bibr" rid="B35">Hackett et al. 2008</xref>) and subsequent synthesis centers, sharing and reuse through synthesis became an established norm for disciplines like ecology, evolution, and socio-ecology. Nonetheless, only recently has data citation been common in synthesis papers. It appears a culture of data citation must be preceded by a culture of sharing and reuse; one that values reproducibility, transparency, and credit in practice (<xref ref-type="bibr" rid="B66">Stuart 2017</xref>).</p>
</sec>
<sec>
<title>III. Current Activity and Issues</title>
<p>In the last few years, we have seen much activity to promote data citation and to define specific guidelines for both data and software citation. The Digital Curation Centre and Earth Science Information Partners (ESIP) have updated their respective, long-standing data citation guidelines (<xref ref-type="bibr" rid="B6">Ball &amp; Duke 2015</xref>; <xref ref-type="bibr" rid="B29">EDPSC 2019</xref>). The Research Data Alliance (RDA) produced a Recommendation on citing specific subsets of very dynamic data (<xref ref-type="bibr" rid="B60">Rauber et al. 2015</xref>). Publishers and repositories are coming together on common citation practices (<xref ref-type="bibr" rid="B22">Cousijn et al. 2018</xref>; <xref ref-type="bibr" rid="B31">Fenner et al. 2019</xref>). Force11 and ESIP have collaborated on software citation principles and guidelines (<xref ref-type="bibr" rid="B30">ESSCC 2019</xref>; <xref ref-type="bibr" rid="B42">Katz &amp; Chue Hong 2018</xref>; <xref ref-type="bibr" rid="B63">Smith et al. 2016</xref>).</p>
<p>We are optimistic that data (and software) citation is emerging as a norm for observational science. We are especially encouraged by the recent project led by the American Geophysical Union and others on &#8216;Enabling FAIR Data&#8217; which has led to many publishers now requiring data citation in their author guidelines (<xref ref-type="bibr" rid="B64">Stall et al. 2018</xref>). The new Transparency and Openness Promotion statement shows similar commitment beyond Earth, environmental, and space sciences (<xref ref-type="bibr" rid="B1">Aalbersberg et al. 2018</xref>). It appears we are approaching the critical mass for a broad behavioral shift. Nonetheless, multiple issues remain.</p>
<p>Many of the issues are rooted in the fact that we are taking a concept implemented for physically printed literature and human beings and trying to use it to address multiple concerns for both humans and machines. We awkwardly try to have the digital space match the &#8216;paper&#8217; space, as Paskin (<xref ref-type="bibr" rid="B58">2000</xref>) put it.</p>
<sec>
<title>A. Specific and verifiable citation</title>
<p>One issue of data citation practice is citing precise subsets of versioned data. This is necessary to meet the &#8216;specific and verifiable&#8217; principle of the <italic>Joint Declaration</italic> and is fundamental to the reproducibility use case for citation. Of course, the simplest, logical approach is to assign a new PID if there is any change in the data set, but this can become unwieldy with large, dynamic data such as data streaming from a remote instrument undergoing multiple calibrations and corrections. Furthermore, repositories can have very different approaches to versioning their data and how they recommend citing different versions. They also package data and assign PIDs at very different levels of granularity.</p>
<p>We find the RDA Recommendation on Data Citation of Evolving Data (<xref ref-type="bibr" rid="B60">Rauber et al. 2015</xref>) the best approach to citing specific subsets of very dynamic data. The basic idea is to assign and maintain a PID for a specific, time-stamped query of a data set, as well as a PID for the data set as a whole. This means the repository must continue to resolve the query PID and maintain or migrate the technology necessary to resolve the actual query within the data set. The RDA Data Citation WG has conducted multiple implementation workshops and has reports from repositories adopting this approach every six months at RDA Plenaries. The approach is getting broader adoption, but it is not at all a norm across repositories. Many simply do not have the capacity to implement it yet; sustaining these citations means sustaining query systems not just data; and maintaining access to these cited queries through technology cycles can be quite challenging (<xref ref-type="bibr" rid="B65">Stockhause &amp; Lautenschlager 2017</xref>). Note, this approach provides reference and access to a precise subset, but it does not necessarily address specific credit concerns for that subset, such as when different authors contribute to a larger collection.</p>
<p>There are other approaches for citing dynamic data recommended by DCC, ESIP, and DataVerse (<xref ref-type="bibr" rid="B6">Ball &amp; Duke 2015</xref>; <xref ref-type="bibr" rid="B24">Crosas 2014</xref>; <xref ref-type="bibr" rid="B29">EDPSC 2019</xref>). These include capturing time slices or snapshots of an ongoing time series, having different PIDs for the data set concept and specific versions or instances, or simply establishing careful documentation practices. These approaches are incomplete in that they tend to be more appropriate for relatively static data and often require human interpretation. For example, it is not practical to continually mint new PIDs for a data set that may update every six seconds (typical for many automated meteorological stations). Similarly, some repositories do not find it appropriate to mint new PIDs for minor changes to a data set (e.g., a typo in the documentation) because it can unnecessarily complicate tracking the use of a data set. They rely on the user to apply their judgement on what is a meaningful change for their application. Computers cannot exercise such judgement.</p>
</sec>
<sec>
<title>B. What to cite</title>
<p>Another issue is deciding what constitutes a first-class object in scholarly discourse &#8212; the &#8216;importance&#8217; principle. Data are only truly useful if they are accompanied with detailed documentation about how they were collected, their uncertainties, and relevant applications. Multiple data journals have emerged to provide &#8216;peer-review&#8217; of data and documentation and to publish &#8216;data papers&#8217; that provide recognizable credit for data authors or creators. They have different approaches, however, on what is to be cited&#8212;the document, the data, or both. Often the paper and the associated data set have different authors. The papers also differ in how they structure the information for humans and machines.</p>
<p>Some organizations are now developing well-structured, machine-readable &#8216;publications&#8217; that provide the best services of both publishers and data repositories. A good example is the Whole Tale project which is a collaboration among repositories (e.g., DataOne, Globus, DataVerse), computing providers, and publishers to implement reproducible data papers (<xref ref-type="bibr" rid="B15">Brinckman et al. 2019</xref>; <xref ref-type="bibr" rid="B19">Chard et al. 2019</xref>). Whole Tale and similar systems like Binder<xref ref-type="fn" rid="n4">4</xref> and CodeOcean<xref ref-type="fn" rid="n5">5</xref> provide mechanisms to package the data and products of research along with the code and computing environment that produced them, machine-readable provenance about how they were produced, references to published input data, and the scientific narrative that frames the rationale and conclusions for the work, all in a citable and re-executable Research Object (<xref ref-type="bibr" rid="B7">Bechhofer et al. 2010</xref>). These complex, hybrid publications span the boundaries between data, software, and publications to enable fully transparent research publication.</p>
</sec>
<sec>
<title>C. Tracking use and impact</title>
<p>Part of the purpose of citing data is to provide credit and attribution for the creation of the data set and correspondingly to help determine how and how often a data set is used. Data can have many important uses outside of research publications, though, and people are exploring better ways to track the impact of data. The National Information Standards Organization (NISO) defined a &#8216;Recommended Practice&#8217; for alternative assessment metrics, but they primarily emphasize the need for data citation. They observe that &#8216;there currently seems to be a lack of interest in altmetrics for data in the community&#8217; (<xref ref-type="bibr" rid="B54">NISO 2016: 16</xref>). Peters et al. (<xref ref-type="bibr" rid="B59">2016</xref>) also find little use of altmetrics and no correlation between altmetrics and citation. Because of inconsistent citation practices, text mining may be a better way to identify literature-data relationships (<xref ref-type="bibr" rid="B38">Kafkas et al. 2013</xref>). Nevertheless, research from Kratz and Strasser (<xref ref-type="bibr" rid="B47">2015b</xref>) indicates that citation and data downloads are the measures most valued by researchers. To that end, work through RDA has led to the &#8216;Make Data Count&#8217; project (<xref ref-type="bibr" rid="B23">Cousijn et al. 2019</xref>; <xref ref-type="bibr" rid="B46">Kratz &amp; Strasser 2015a</xref>), which has defined a consistent way to count data downloads through the COUNTER Code of Practice for Research Data (<xref ref-type="bibr" rid="B32">Fenner et al. 2018</xref>). These projects are still primarily oriented to the research community.</p>
<p>Much more research is needed on how to assess data use and impact beyond bibliometrics. Groups such as the VALUABLES consortium, are beginning to explore these issues through the lens of economics. For example, Bernknopf, et al. (<xref ref-type="bibr" rid="B9">2016</xref>) found that federal agencies could save $7.7 million/year in post-wildfire response if Landsat data were used; while Cooke &amp; Golub (<xref ref-type="bibr" rid="B20">2019</xref>) report that a 30% reduction in weather uncertainty impacts on corn and soybean futures due to soil moisture measurements from NASA&#8217;s Soil Moisture Active Passive satellite had a net worth of $1.44 Billion/year. Even more uncertain are methods of quantifying the impacts of data on public policy.</p>
<p>Providing credit for that impact is also tricky. Many people play critical roles in the creation of even the simplest data set, and their performance is evaluated in different ways. Not everyone is measured by their research paper publication record. One effort to recognize these other contributions is Project CRediT (Contributor Roles Taxonomy), which has defined a taxonomy of contributor roles for research objects and suggests using digital badges that detail what each author did for the work and link to their profiles elsewhere on the Web (<xref ref-type="bibr" rid="B2">Allen et al. 2014</xref>). We also recognize the concept of transitive credit, which could be used to recognize the developers of products other than papers (<xref ref-type="bibr" rid="B41">Katz 2014</xref>), and to use provenance to understand how upstream data and software enabled advances in research. Ultimately, we must recognize that credit is a human concern requiring context and interpretation that cannot be readily automated. We can work to build attribution into the scientific workflow, but human judgement is still required to assess the relative value of various contributions. Indeed, a study of software attribution found that automating credit mechanisms can lead to perverse metrics and incentives that can falsely represent the value of a contribution (<xref ref-type="bibr" rid="B3">Alliez et al. 2019</xref>).</p>
<p>Another concern related to both credit and understanding impact is to identify and trace the connections between all sorts of research objects (data, literature, software, people, organizations, algorithms, etc). Multiple efforts try to address this. One effort directly related to scholarly publishing is the Scholix (Scholarly Link Exchange) initiative which emerged out of RDA to more formally interconnect data and literature (<xref ref-type="bibr" rid="B17">Burton et al. 2017</xref>). The approach is functional and operational, as it builds from established citation hubs like DataCite, CrossRef, and OpenAire. This effort is expanding and collaborating with other related initiatives through a newly proposed &#8216;Open Science Graphs for FAIR Data Interest Group&#8217; within RDA.</p>
<p>Other approaches use more decentralized, Web-based mechanisms which may allow more adaptability and extensibility (e.g., <xref ref-type="bibr" rid="B50">Ma et al. 2017</xref>; <xref ref-type="bibr" rid="B57">Parsons &amp; Fox 2018</xref>). Work in ESIP and RDA explores how we can use <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://schema.org">schema.org</ext-link> web markup to identify and link data and repositories. ESIP is also exploring another W3C recommendation called Linked Data Notifications<xref ref-type="fn" rid="n6">6</xref> &#8212; a sort of RSS-style protocol to request and receive notifications about activities, interactions, and new information. The notification itself is an individual entity with its own URI. Another emergent approach is the Digital Object Interface Protocol (<xref ref-type="bibr" rid="B40">Kahn et al. 2018</xref>) which assigns a persistent ID to any digital object and allows that object to express what type of object it is and what operations it allows such as various web services or basic management functions.</p>
<p>All these interconnecting technologies are somewhat peripheral to the core purpose of citation, but they highlight how applying machine-actionable PIDs to digital objects can expand the possibilities for knowledge sharing well beyond the bounds of traditional (paper-based) citation. This brings us to the issue and concern of identity itself.</p>
</sec>
<sec>
<title>D. Identifying things</title>
<p>People are beginning to rethink how PIDs work. To date, the basic issue of persistence of locators on the web has been addressed by what we might call authority-based identifiers, which separate the identity of an object from its location and are maintained in trusted registries. Klump et al. (<xref ref-type="bibr" rid="B44">2015</xref>, <xref ref-type="bibr" rid="B45">2017</xref>) discuss how this approach has evolved and raise issues around the interconnection of identity, institutional commitment, and cost models. They note: &#8216;The focus of the DOI for the data community on paper-like documents and human actors has left some conceptual gaps&#8217; (<xref ref-type="bibr" rid="B44">Klump et al. 2015: 133</xref>). They argue that we need to explore more advanced features such as identifier templates, more sophisticated content negotiation when resolving identifiers, more machine actionability in general, as well as the social process of maintaining the persistence of an object and its reference.</p>
<p>In other work, data managers are looking to content-based identifiers (i.e., cryptographic-hash-based IDs) to identify exact copies of data, ideally without relying on third parties and external administrative processes. These content-based identifiers can be deployed and resolved in peer-to-peer environments like the InterPlanetary File System (IPFS)<xref ref-type="fn" rid="n7">7</xref> and Dat.<xref ref-type="fn" rid="n8">8</xref> There is already an established system called Qri<xref ref-type="fn" rid="n9">9</xref> (query) which allows users to reference, browse, download, create, fork, and publish data sets with a broad network of peers in IPFS. Furthermore, Dat includes public-key technology to provide assurance on the source of the data and any changes that may have occurred.<xref ref-type="fn" rid="n10">10</xref> These approaches are still not well-suited to massive volumes of streaming data, and there is still the issue that different representations of data may be scientifically equivalent but not identical. The hash really needs to include the provenance chain as well as the data set. Nonetheless, these approaches show great promise. In related work, Bolikowski et al. (<xref ref-type="bibr" rid="B12">2015: 281</xref>) use the concepts of blockchain and version control systems like Git to propose a distributed system for maintaining long-term resolvability of persistent identifiers. They argue that the &#8216;system should be agnostic with respect to referent type (data sets, source codes, documents, people) and content delivery technology (HTTP, BitTorrent, Tor/Onion)&#8217;.</p>
<p>It is important to note that content-based identifiers have quite different properties and applications from authority-based identifiers. Content-based identifiers and blockchain technologies can be useful for tracking provenance, precise data queries, and internal repository management concerns. Authority-based identifiers are being used for ensuring the social requirements necessary to maintain a persistent and managed access location (<xref ref-type="bibr" rid="B26">Di Cosmo et al. 2018</xref>), while those governance structures are only now being developed for distributed, content-based identifier systems.</p>
<p>Finally, it is worth noting that various authority-based identifiers have or are being developed for other research objects and entities. These include the Open Researcher and Contributor ID (ORCID) for individual researchers (<xref ref-type="bibr" rid="B34">Haak et al. 2012</xref>); the Research Organization Registry (ROR) Community, which is working to develop identifiers for research organizations;<xref ref-type="fn" rid="n11">11</xref> and the work of the Persistent Identification of Instruments Working Group of RDA,<xref ref-type="fn" rid="n12">12</xref> which is implementing processes to use DataCite DOIs for scientific instruments. These are only a few examples of the growth in the development and application of PIDs. Similar to Linked Open Data approaches (<xref ref-type="bibr" rid="B8">Bechhofer et al. 2011</xref>; <xref ref-type="bibr" rid="B11">Bizer et al. 2009</xref>), these other-types of PIDs can make more elements of a citation precise and machine-actionable, but again they move us further away from the traditional human-oriented citation. Consider, fancifully, if this article was cited making full use of identifiers (Figure <xref ref-type="fig" rid="F1">1</xref>). It is more precise but also more opaque for the human reader.</p>
<fig id="F1">
<label>Figure 1</label>
<caption>
<p>Imaginary citation for this article making full use of PIDs.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="dsj-18-1048-g1.png"/>
</fig>
<p>Also, PIDs cannot always point precisely to the thing but rather only a representation of the thing (ORCIDs don&#8217;t reference people, they reference descriptions of people). Access may be precisely defined, but credit and reference are inherently ambiguous to some degree (<xref ref-type="bibr" rid="B36">Hayes &amp; Halpin 2008</xref>). The human remains in the loop.</p>
</sec>
</sec>
<sec>
<title>IV. Looking to the Future</title>
<p>Despite recognized definitions (<xref ref-type="bibr" rid="B13">Borgman 2015</xref>; <xref ref-type="bibr" rid="B25">DCSG 2014</xref>; <xref ref-type="bibr" rid="B67">TGDCSP 2013</xref>), data citation remains a complex and evolving issue. On the one hand, the basic principles and process are well established. We know how to cite most data in research publications. We must only accelerate the implementation, and there does appear to be movement in that direction. On the other hand, long-established academic practices and assumptions about what is &#8216;important&#8217; in scientific work lead us to bundle many different concerns into the concept of citation. At the same time, the opportunities promised by PIDs lead us to bundle even more concerns into reference schemes. Reconsideration and disaggregation of data citation concerns are overdue.</p>
<p>Data citation is often viewed as a computational or information science problem (<xref ref-type="bibr" rid="B16">Buneman et al. 2016</xref>; <xref ref-type="bibr" rid="B62">Silvello 2018</xref>), and sometimes more accurately as a social or cultural adaptation problem (<xref ref-type="bibr" rid="B13">Borgman 2015</xref>; <xref ref-type="bibr" rid="B44">Klump et al. 2015</xref>). But it is a complex socio-technical problem with many nuanced concerns. In short, it is an issue of praxis.</p>
<p>We are inspired by the Force11 effort to identify some of the myriad use cases for software citation (<xref ref-type="bibr" rid="B63">Smith et al. 2016</xref>). We feel we need to do the same with data citation: define multiple use cases that 1) de-emphasize the importance of the scientific paper in lieu of more precise assertions and supporting evidence and 2) emphasize the valuable use of data outside traditional scholarly environments. Some of this work has begun in RDA and ESIP. This should help us sort out the different concerns.</p>
<p>We already know that credit is primarily a human concern and access is a machine concern (reference is both), but what does that mean in practice? In health and social sciences, researchers have developed a &#8216;Payback Framework&#8217; with a logical model of the complete research process and categories of (social health) payback from research (<xref ref-type="bibr" rid="B27">Donovan and Hanney 2011</xref>). Can we extend this and apply it to data by recognizing the reuse and value generated at many different stages? Can machine-actionable badges capture credit better than centralized citation indices?</p>
<p>Recognizing access as a machine concern can help us focus on providing data as a service rather than simply as object downloaded by a human. This, in turn, can help us make intelligent choices about what type of identifiers to use for what application. The ID for the human interested in a general description of the data may be different and will behave differently than the ID for the machine.</p>
<p>Going forward, we should accelerate the substantial progress we have made on implementing data citation for the basic scholarly use case. At the same time, we should not overextend the concept nor expand our expectations for what citation can accomplish. It is time to rethink some of our assumptions if we are to make data both FAIR and fair.</p>
</sec>
</body>
<back>
<fn-group>
<fn id="n1"><p>DOI Factsheet at <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.doi.org/factsheets/DOIKeyFacts.html">https://www.doi.org/factsheets/DOIKeyFacts.html</ext-link>.</p></fn>
<fn id="n2"><p><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://n2t.net/e/ark_ids.html">https://n2t.net/e/ark_ids.html</ext-link>.</p></fn>
<fn id="n3"><p><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://n2t.net">https://n2t.net</ext-link>.</p></fn>
<fn id="n4"><p><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://mybinder.org">https://mybinder.org</ext-link>.</p></fn>
<fn id="n5"><p><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://codeocean.com">https://codeocean.com</ext-link>.</p></fn>
<fn id="n6"><p>Linked Data Notifications: W3C Recommendation 2 May 2017 <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.w3.org/TR/2017/REC-ldn-20170502/">https://www.w3.org/TR/2017/REC-ldn-20170502/</ext-link>.</p></fn>
<fn id="n7"><p><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://ipfs.io">https://ipfs.io</ext-link>.</p></fn>
<fn id="n8"><p><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://dat.foundation">https://dat.foundation</ext-link>.</p></fn>
<fn id="n9"><p><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://qri.io">https://qri.io</ext-link>.</p></fn>
<fn id="n10"><p><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://datprotocol.github.io/how-dat-works/">https://datprotocol.github.io/how-dat-works/</ext-link>.</p></fn>
<fn id="n11"><p><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.ror.community">https://www.ror.community</ext-link>.</p></fn>
<fn id="n12"><p><ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.rd-alliance.org/groups/persistent-identification-instruments-wg">https://www.rd-alliance.org/groups/persistent-identification-instruments-wg</ext-link>.</p></fn>
</fn-group>
<ack>
<title>Acknowledgements</title>
<p>This article is partially based on work done with community support provided by the Research Data Alliance and the Earth Science Information Partners. Parsons was partially supported by award G-2018-11204 from the AP Sloan Foundation. Duerr was partially supported by National Science Foundation award #1639753. Jones was partially supported by the National Science Foundation (awards #1546024 and #1430508) and the National Center for Ecological Analysis and Synthesis, a Center funded by the University of California, Santa Barbara, and the State of California.</p>
</ack>
<sec>
<title>Competing Interests</title>
<p>All the authors have been active in the development of data citation guidelines and implementations. This includes participating in relevant task groups and committees of multiple organizations, including but not limited to CODATA, DataOne, ESIP, Force11, and RDA. They have implemented data citation practice at multiple repositories, including but limited to the Arctic Data Center, the Deep Carbon Observatory Data Portal, the KNB Data Repository, and the National Snow and Ice Data Center. They co-authored data citation and usage standards including DCSG (<xref ref-type="bibr" rid="B25">2014</xref>), EDPSC (<xref ref-type="bibr" rid="B29">2019</xref>), Fenner et al. (<xref ref-type="bibr" rid="B32">2018</xref>), and TGDCSP (<xref ref-type="bibr" rid="B67">2013</xref>). MP is the Editor in Chief of the Data Science Journal, but did not oversee the review of this article.</p>
</sec>
<ref-list>
<ref id="B1"><label>1</label><mixed-citation publication-type="journal"><string-name><surname>Aalbersberg</surname>, <given-names>IJ</given-names></string-name>, et al. <year>2018</year>. <article-title>Making science transparent by default; introducing the TOP statement</article-title>. DOI: <pub-id pub-id-type="doi">10.31219/osf.io/sm78t</pub-id></mixed-citation></ref>
<ref id="B2"><label>2</label><mixed-citation publication-type="journal"><string-name><surname>Allen</surname>, <given-names>L</given-names></string-name>, et al. <year>2014</year>. <article-title>Publishing: Credit where credit is due</article-title>. <source>Nature</source>, <volume>508</volume>: <fpage>312</fpage>&#8211;<lpage>313</lpage>. DOI: <pub-id pub-id-type="doi">10.1038/508312a</pub-id></mixed-citation></ref>
<ref id="B3"><label>3</label><mixed-citation publication-type="webpage"><string-name><surname>Alliez</surname>, <given-names>P</given-names></string-name>, et al. <year>2019</year>. <article-title>Attributing and referencing (research) software: Best practices and outlook from Inria</article-title>. <source>Computing in Science &amp; Engineering</source>. <uri>https://arxiv.org/abs/1905.11123</uri>.</mixed-citation></ref>
<ref id="B4"><label>4</label><mixed-citation publication-type="webpage"><string-name><surname>Altman</surname>, <given-names>M</given-names></string-name> and <string-name><surname>King</surname>, <given-names>G</given-names></string-name>. <year>2007</year>. <article-title>A proposed standard for the scholarly citation of quantitative data</article-title>. <source>D-Lib Magazine</source>, <fpage>13</fpage>. <uri>http://dlib.org/dlib/march07/altman/03altman.html</uri> accessed 2019-07-25.</mixed-citation></ref>
<ref id="B5"><label>5</label><mixed-citation publication-type="journal"><string-name><surname>Baker</surname>, <given-names>KS</given-names></string-name> and <string-name><surname>Yarmey</surname>, <given-names>L</given-names></string-name>. <year>2009</year>. <article-title>Data stewardship: Environmental Data Curation and a Web-of-Repositories</article-title>. <source>International Journal of Digital Curation</source>, <fpage>4</fpage>. DOI: <pub-id pub-id-type="doi">10.2218/ijdc.v4i2.90</pub-id></mixed-citation></ref>
<ref id="B6"><label>6</label><mixed-citation publication-type="webpage"><string-name><surname>Ball</surname>, <given-names>A</given-names></string-name> and <string-name><surname>Duke</surname>, <given-names>M</given-names></string-name>. <year>2015</year>. <chapter-title>How to Cite Datasets and Link to Publications</chapter-title>. <publisher-loc>Edinburgh</publisher-loc>: <publisher-name>Digital Curation Centre</publisher-name>. <uri>http://www.dcc.ac.uk/resources/how-guides</uri> accessed 2019-02-06.</mixed-citation></ref>
<ref id="B7"><label>7</label><mixed-citation publication-type="webpage"><string-name><surname>Bechhofer</surname>, <given-names>S</given-names></string-name>, et al. <year>2010</year>. <article-title>Research objects: Towards exchange and reuse of digital knowledge</article-title>. <source>The Future of the Web for Collaborative Science (FWCS 2010)</source>. <uri>https://eprints.soton.ac.uk/268555/</uri> accessed 2019-07-28.</mixed-citation></ref>
<ref id="B8"><label>8</label><mixed-citation publication-type="journal"><string-name><surname>Bechhofer</surname>, <given-names>S</given-names></string-name>, et al. <year>2011</year>. <article-title>Why linked data is not enough for scientists</article-title>. <source>Future Generation Computer Systems</source>. DOI: <pub-id pub-id-type="doi">10.1016/j.future.2011.08.004</pub-id></mixed-citation></ref>
<ref id="B9"><label>9</label><mixed-citation publication-type="webpage"><string-name><surname>Bernknopf</surname>, <given-names>R</given-names></string-name>, et al. <year>2016</year>. <article-title>The cost-effectiveness of satellite Earth observations to inform a post-wildfire response</article-title>. <source>Working Paper</source>, <fpage>19</fpage>&#8211;<lpage>16</lpage>. <uri>https://media.rff.org/documents/Valuables_Wildfires.pdf</uri> accessed 2019-07-28.</mixed-citation></ref>
<ref id="B10"><label>10</label><mixed-citation publication-type="webpage"><string-name><surname>Berquist</surname>, <given-names>CR</given-names>, <suffix>Jr</suffix></string-name>. <year>1999</year>. <article-title>Digital map production and publication by geological survey organizations: A proposal for authorship and citation guidelines</article-title>. <source>U.S. Geological Survey Open-File Report</source>, <fpage>99</fpage>&#8211;<lpage>386</lpage>. <uri>https://pubs.usgs.gov/of/1999/of99-386/berquist.html</uri> accessed 2019-03-02.</mixed-citation></ref>
<ref id="B11"><label>11</label><mixed-citation publication-type="journal"><string-name><surname>Bizer</surname>, <given-names>C</given-names></string-name>, <string-name><surname>Heath</surname>, <given-names>T</given-names></string-name> and <string-name><surname>Berners-Lee</surname>, <given-names>T</given-names></string-name>. <year>2009</year>. <article-title>Linked data &#8211; the story so far</article-title>. <source>International Journal on Semantic Web and Information Systems</source>, <volume>5</volume>: <fpage>1</fpage>&#8211;<lpage>22</lpage>. DOI: <pub-id pub-id-type="doi">10.4018/jswis.2009081901</pub-id></mixed-citation></ref>
<ref id="B12"><label>12</label><mixed-citation publication-type="journal"><string-name><surname>Bolikowski</surname>, <given-names>L</given-names></string-name>, <string-name><surname>Nowi&#324;ski</surname>, <given-names>A</given-names></string-name> and <string-name><surname>Sylwestrzak</surname>, <given-names>W</given-names></string-name>. <year>2015</year>. <article-title>A system for distributed minting and management of persistent identifiers</article-title>. <source>International Journal of Digital Curation</source>, <volume>10</volume>: <fpage>280</fpage>&#8211;<lpage>286</lpage>. DOI: <pub-id pub-id-type="doi">10.2218/ijdc.v10i1.368</pub-id></mixed-citation></ref>
<ref id="B13"><label>13</label><mixed-citation publication-type="book"><string-name><surname>Borgman</surname>, <given-names>C</given-names></string-name>. <year>2015</year>. <source>Big Data, Little Data, No Data</source>. <publisher-loc>Boston</publisher-loc>: <publisher-name>MIT Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.7551/mitpress/9963.001.0001</pub-id></mixed-citation></ref>
<ref id="B14"><label>14</label><mixed-citation publication-type="webpage"><string-name><surname>Borgman</surname>, <given-names>C</given-names></string-name>. <year>2016</year>. <chapter-title>Data citation as a bibliometric oxymoron</chapter-title>. In: <source>Theories of Informetrics and Scholarly Communication</source>, <string-name><surname>Sugimoto</surname>, <given-names>CR</given-names></string-name> (ed.), <fpage>93</fpage>&#8211;<lpage>115</lpage>. <publisher-loc>Berlin &amp; Boston</publisher-loc>: <publisher-name>Walter de Gruyter GmbH &amp; Co KG</publisher-name>. <uri>https://escholarship.org/content/qt8w36p9zf/qt8w36p9zf.pdf</uri> accessed 2019-07-26.</mixed-citation></ref>
<ref id="B15"><label>15</label><mixed-citation publication-type="journal"><string-name><surname>Brinckman</surname>, <given-names>A</given-names></string-name>, et al. <year>2019</year>. <article-title>Computing environments for reproducibility: Capturing the &#8220;whole tale&#8221;</article-title>. <source>Future Generation Computer Systems</source>, <volume>94</volume>: <fpage>854</fpage>&#8211;<lpage>867</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.future.2017.12.029</pub-id></mixed-citation></ref>
<ref id="B16"><label>16</label><mixed-citation publication-type="journal"><string-name><surname>Buneman</surname>, <given-names>P</given-names></string-name>, <string-name><surname>Davidson</surname>, <given-names>S</given-names></string-name> and <string-name><surname>Frew</surname>, <given-names>J</given-names></string-name>. <year>2016</year>. <article-title>Why data citation is a computational problem</article-title>. <source>Commun ACM</source>, <volume>59</volume>: <fpage>50</fpage>&#8211;<lpage>57</lpage>. DOI: <pub-id pub-id-type="doi">10.1145/2893181</pub-id></mixed-citation></ref>
<ref id="B17"><label>17</label><mixed-citation publication-type="journal"><string-name><surname>Burton</surname>, <given-names>A</given-names></string-name>, et al. <year>2017</year>. <article-title>The Scholix framework for interoperability in data-literature information exchange</article-title>. <source>D-Lib Magazine</source>, <fpage>23</fpage>. DOI: <pub-id pub-id-type="doi">10.1045/january2017-burton</pub-id></mixed-citation></ref>
<ref id="B18"><label>18</label><mixed-citation publication-type="webpage"><string-name><surname>Callaghan</surname>, <given-names>S</given-names></string-name>, et al. <year>2009</year>. <article-title>Overlay journals and data publishing in the meteorological sciences</article-title>. <source>Ariadne</source>. <uri>http://www.ariadne.ac.uk/issue60/callaghan-et-al/</uri> accessed 2011-11-27.</mixed-citation></ref>
<ref id="B19"><label>19</label><mixed-citation publication-type="journal"><string-name><surname>Chard</surname>, <given-names>K</given-names></string-name>, et al. <year>2019</year>. <article-title>Implementing computational reproducibility in the Whole Tale environment</article-title>. <source>Proceedings of the 2nd International Workshop on Practical Reproducible Evaluation of Computer Systems &#8211; P-RECS &#8216;19</source>. DOI: <pub-id pub-id-type="doi">10.1145/3322790.3330594</pub-id></mixed-citation></ref>
<ref id="B20"><label>20</label><mixed-citation publication-type="webpage"><string-name><surname>Cooke</surname>, <given-names>R</given-names></string-name> and <string-name><surname>Golub</surname>, <given-names>A</given-names></string-name>. <year>2019</year>. <article-title>Market-based methods for monetizing uncertainty reduction: A case study</article-title>. <source>Working Paper</source>, <fpage>19</fpage>&#8211;<lpage>15</lpage>. <uri>https://media.rff.org/documents/WP_Cooke_Golub_4.pdf</uri> accessed 2019-07-28.</mixed-citation></ref>
<ref id="B21"><label>21</label><mixed-citation publication-type="journal"><string-name><surname>Costello</surname>, <given-names>MJ</given-names></string-name>. <year>2009</year>. <article-title>Motivating online publication of data</article-title>. <source>Bioscience</source>, <volume>59</volume>: <fpage>418</fpage>&#8211;<lpage>427</lpage>. DOI: <pub-id pub-id-type="doi">10.1525/bio.2009.59.5.9</pub-id></mixed-citation></ref>
<ref id="B22"><label>22</label><mixed-citation publication-type="journal"><string-name><surname>Cousijn</surname>, <given-names>H</given-names></string-name>, et al. <year>2018</year>. <article-title>A data citation roadmap for scientific publishers</article-title>. <source>Sci Data</source>, <volume>5</volume>: <elocation-id>180259</elocation-id>. DOI: <pub-id pub-id-type="doi">10.1038/sdata.2018.259</pub-id></mixed-citation></ref>
<ref id="B23"><label>23</label><mixed-citation publication-type="journal"><string-name><surname>Cousijn</surname>, <given-names>H</given-names></string-name>, et al. <year>2019</year>. <article-title>Bringing citations and usage metrics together to make data count</article-title>. <source>Data Science Journal</source>, <fpage>18</fpage>. DOI: <pub-id pub-id-type="doi">10.5334/dsj-2019-009</pub-id></mixed-citation></ref>
<ref id="B24"><label>24</label><mixed-citation publication-type="journal"><string-name><surname>Crosas</surname>, <given-names>M</given-names></string-name>. <year>2014</year>. <article-title>The evolution of data citation: From principles to implementation</article-title>. <source>IASSIST Quarterly</source>, <volume>37</volume>: <fpage>62</fpage>. DOI: <pub-id pub-id-type="doi">10.29173/iq504</pub-id></mixed-citation></ref>
<ref id="B25"><label>25</label><mixed-citation publication-type="journal"><collab>DCSG &#8211; Data Citation Synthesis Group</collab>. <year>2014</year>. <source>Joint Declaration of Data Citation Principles</source>. DOI: <pub-id pub-id-type="doi">10.25490/a97f-egyk</pub-id></mixed-citation></ref>
<ref id="B26"><label>26</label><mixed-citation publication-type="journal"><string-name><surname>Di Cosmo</surname>, <given-names>R</given-names></string-name>, <string-name><surname>Gruenpeter</surname>, <given-names>M</given-names></string-name> and <string-name><surname>Zacchiroli</surname>, <given-names>S</given-names></string-name>. <year>2018</year>. <article-title>Identifiers for digital objects: The case of software source code preservation</article-title>. <source>Open Science Framework</source>. DOI: <pub-id pub-id-type="doi">10.17605/OSF.IO/KDE56</pub-id></mixed-citation></ref>
<ref id="B27"><label>27</label><mixed-citation publication-type="journal"><string-name><surname>Donovan</surname>, <given-names>C</given-names></string-name> and <string-name><surname>Hanney</surname>, <given-names>S</given-names></string-name>. <year>2011</year>. <article-title>The payback framework explained</article-title>. <source>Research Evaluation</source>, <volume>20</volume>: <fpage>181</fpage>&#8211;<lpage>183</lpage>. DOI: <pub-id pub-id-type="doi">10.3152/095820211X13118583635756</pub-id></mixed-citation></ref>
<ref id="B28"><label>28</label><mixed-citation publication-type="journal"><string-name><surname>Downs</surname>, <given-names>RR</given-names></string-name>, <string-name><surname>Duerr</surname>, <given-names>R</given-names></string-name>, <string-name><surname>Hills</surname>, <given-names>DJ</given-names></string-name> and <string-name><surname>Ramapriyan</surname>, <given-names>HK</given-names></string-name>. <year>2015</year>. <article-title>Data stewardship in the Earth sciences</article-title>. <source>D-Lib Magazine</source>, <fpage>21</fpage>. DOI: <pub-id pub-id-type="doi">10.1045/july2015-downs</pub-id></mixed-citation></ref>
<ref id="B29"><label>29</label><mixed-citation publication-type="book"><collab>EDPSC &#8211; ESIP Data Preservation and Stewardship Committee</collab>. <year>2019</year>. <source>Data Citation Guidelines for Earth Science Data, Version 2</source>. <publisher-name>Earth Science Information Partners</publisher-name>. DOI: <pub-id pub-id-type="doi">10.6084/m9.figshare.8441816.v1</pub-id></mixed-citation></ref>
<ref id="B30"><label>30</label><mixed-citation publication-type="book"><collab>ESSCC &#8211; ESIP Software and Services Citation Cluster</collab>. <year>2019</year>. <source>Software and Services Citation Guidelines and Examples. Ver. 1</source>. <publisher-name>Earth Science Information Partners</publisher-name>. DOI: <pub-id pub-id-type="doi">10.6084/m9.figshare.7640426</pub-id></mixed-citation></ref>
<ref id="B31"><label>31</label><mixed-citation publication-type="journal"><string-name><surname>Fenner</surname>, <given-names>M</given-names></string-name>, et al. <year>2019</year>. <article-title>A data citation roadmap for scholarly data repositories</article-title>. <source>Scientific Data</source>, <volume>6</volume><issue>(1) (1)</issue>: <fpage>28</fpage>. DOI: <pub-id pub-id-type="doi">10.1038/s41597-019-0031-8</pub-id></mixed-citation></ref>
<ref id="B32"><label>32</label><mixed-citation publication-type="journal"><string-name><surname>Fenner</surname>, <given-names>M</given-names></string-name>, et al. <year>2018</year>. <article-title>Code of practice for research data usage metrics release 1</article-title>. DOI: <pub-id pub-id-type="doi">10.7287/peerj.preprints.26505v1</pub-id></mixed-citation></ref>
<ref id="B33"><label>33</label><mixed-citation publication-type="journal"><string-name><surname>Guralnick</surname>, <given-names>RP</given-names></string-name>, et al. <year>2015</year>. <article-title>Community next steps for making globally unique identifiers work for biocollections data</article-title>. <source>Zookeys</source>, <fpage>133</fpage>&#8211;<lpage>154</lpage>. DOI: <pub-id pub-id-type="doi">10.3897/zookeys.494.9352</pub-id></mixed-citation></ref>
<ref id="B34"><label>34</label><mixed-citation publication-type="journal"><string-name><surname>Haak</surname>, <given-names>LL</given-names></string-name>, et al. <year>2012</year>. <article-title>Orcid: A system to uniquely identify researchers</article-title>. <source>Learned Publishing</source>, <volume>25</volume>: <fpage>259</fpage>&#8211;<lpage>264</lpage>. DOI: <pub-id pub-id-type="doi">10.1087/20120404</pub-id></mixed-citation></ref>
<ref id="B35"><label>35</label><mixed-citation publication-type="book"><string-name><surname>Hackett</surname>, <given-names>EJ</given-names></string-name>, et al. <year>2008</year>. <chapter-title>Ecology transformed: The national center for ecological analysis and synthesis and the changing patterns of ecological research</chapter-title>. In: <source>Scientific Collaboration on the Internet</source>, <fpage>277</fpage>&#8211;<lpage>296</lpage>. <publisher-name>The MIT Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.7551/mitpress/9780262151207.003.0016</pub-id></mixed-citation></ref>
<ref id="B36"><label>36</label><mixed-citation publication-type="journal"><string-name><surname>Hayes</surname>, <given-names>PJ</given-names></string-name> and <string-name><surname>Halpin</surname>, <given-names>H</given-names></string-name>. <year>2008</year>. <article-title>In defense of ambiguity</article-title>. <source>International Journal on Semantic Web and Information Systems</source>, <volume>4</volume>: <fpage>1</fpage>&#8211;<lpage>18</lpage>. DOI: <pub-id pub-id-type="doi">10.4018/jswis.2008040101</pub-id></mixed-citation></ref>
<ref id="B37"><label>37</label><mixed-citation publication-type="journal"><string-name><surname>Howison</surname>, <given-names>J</given-names></string-name> and <string-name><surname>Bullard</surname>, <given-names>J</given-names></string-name>. <year>2016</year>. <article-title>Software in the scientific literature: Problems with seeing, finding, and using software mentioned in the biology literature</article-title>. <source>Journal of the Association for Information Science and Technology</source>, <volume>67</volume>: <fpage>2137</fpage>&#8211;<lpage>2155</lpage>. DOI: <pub-id pub-id-type="doi">10.1002/asi.23538</pub-id></mixed-citation></ref>
<ref id="B38"><label>38</label><mixed-citation publication-type="journal"><string-name><surname>Kafkas</surname>, <given-names>&#350;</given-names></string-name>, <string-name><surname>Kim</surname>, <given-names>JH</given-names></string-name> and <string-name><surname>McEntyre</surname>, <given-names>JR</given-names></string-name>. <year>2013</year>. <article-title>Database citation in full text biomedical articles</article-title>. <source>PLoS One</source>, <volume>8</volume>: <fpage>e63184</fpage>. DOI: <pub-id pub-id-type="doi">10.1371/journal.pone.0063184</pub-id></mixed-citation></ref>
<ref id="B39"><label>39</label><mixed-citation publication-type="webpage"><string-name><surname>Kahn</surname>, <given-names>R</given-names></string-name> and <string-name><surname>Wilensky</surname>, <given-names>R</given-names></string-name>. <year>1995</year>. <article-title>A framework for distributed digital object services</article-title>. <uri>http://handle.net/cnri.dlib/tn95-01</uri> accessed 2019-07-20.</mixed-citation></ref>
<ref id="B40"><label>40</label><mixed-citation publication-type="webpage"><string-name><surname>Kahn</surname>, <given-names>RE</given-names></string-name>, et al. <year>2018</year>. <source>Digital Object Interface Protocol Specification, Ver. 2.0</source>. <publisher-name>DONA</publisher-name>. <uri>https://www.dona.net/sites/default/files/2018-11/DOIPv2Spec_1.pdf</uri> accessed 2019-07-25.</mixed-citation></ref>
<ref id="B41"><label>41</label><mixed-citation publication-type="journal"><string-name><surname>Katz</surname>, <given-names>DS</given-names></string-name>. <year>2014</year>. <article-title>Transitive credit as a means to address social and technological concerns stemming from citation and attribution of digital products</article-title>. <source>Journal of Open Research Software</source>, <volume>2</volume>: <fpage>e20</fpage>. DOI: <pub-id pub-id-type="doi">10.5334/jors.be</pub-id></mixed-citation></ref>
<ref id="B42"><label>42</label><mixed-citation publication-type="webpage"><string-name><surname>Katz</surname>, <given-names>DS</given-names></string-name> and <string-name><surname>Chue Hong</surname>, <given-names>NP</given-names></string-name>. <year>2018</year>. <article-title>Software citation in theory and practice</article-title>. <source>Arxiv preprint</source>. <uri>https://arxiv.org/pdf/1807.08149.pdf</uri> accessed 2018-12-06.</mixed-citation></ref>
<ref id="B43"><label>43</label><mixed-citation publication-type="journal"><string-name><surname>Klump</surname>, <given-names>J</given-names></string-name>, et al. <year>2006</year>. <article-title>Data publication in the open access initiative</article-title>. <source>Data Science Journal</source>, <volume>5</volume>: <fpage>79</fpage>&#8211;<lpage>83</lpage>. DOI: <pub-id pub-id-type="doi">10.2481/dsj.5.79</pub-id></mixed-citation></ref>
<ref id="B44"><label>44</label><mixed-citation publication-type="journal"><string-name><surname>Klump</surname>, <given-names>J</given-names></string-name>, <string-name><surname>Huber</surname>, <given-names>R</given-names></string-name> and <string-name><surname>Diepenbroek</surname>, <given-names>M</given-names></string-name>. <year>2015</year>. <article-title>DOI for geoscience data-how early practices shape present perceptions</article-title>. <source>Earth Science Informatics</source>, <fpage>1</fpage>&#8211;<lpage>14</lpage>. DOI: <pub-id pub-id-type="doi">10.1007/s12145-015-0231-5</pub-id></mixed-citation></ref>
<ref id="B45"><label>45</label><mixed-citation publication-type="journal"><string-name><surname>Klump</surname>, <given-names>J</given-names></string-name>, <string-name><surname>Murphy</surname>, <given-names>F</given-names></string-name>, <string-name><surname>Weigel</surname>, <given-names>T</given-names></string-name> and <string-name><surname>Parsons</surname>, <given-names>MA</given-names></string-name>. <year>2017</year>. <article-title>20 years of persistent identifiers&#8211;applications and future directions</article-title>. <source>Data Science Journal</source>, <fpage>16</fpage>. DOI: <pub-id pub-id-type="doi">10.5334/dsj-2017-052</pub-id></mixed-citation></ref>
<ref id="B46"><label>46</label><mixed-citation publication-type="journal"><string-name><surname>Kratz</surname>, <given-names>JE</given-names></string-name> and <string-name><surname>Strasser</surname>, <given-names>C</given-names></string-name>. <year>2015a</year>. <article-title>Comment: Making data count</article-title>. <source>Sci Data</source>, <volume>2</volume>: <elocation-id>150039</elocation-id>. DOI: <pub-id pub-id-type="doi">10.1038/sdata.2015.39</pub-id></mixed-citation></ref>
<ref id="B47"><label>47</label><mixed-citation publication-type="journal"><string-name><surname>Kratz</surname>, <given-names>JE</given-names></string-name> and <string-name><surname>Strasser</surname>, <given-names>C</given-names></string-name>. <year>2015b</year>. <article-title>Researcher perspectives on publication and peer review of data</article-title>. <source>PLoS One</source>, <volume>10</volume>: <elocation-id>e0117619</elocation-id>. DOI: <pub-id pub-id-type="doi">10.1371/journal.pone.0117619</pub-id></mixed-citation></ref>
<ref id="B48"><label>48</label><mixed-citation publication-type="journal"><string-name><surname>Lawrence</surname>, <given-names>B</given-names></string-name>, et al. <year>2011</year>. <article-title>Citation and peer review of data: Moving towards formal data publication</article-title>. <source>International Journal of Digital Curation</source>, <fpage>6</fpage>. DOI: <pub-id pub-id-type="doi">10.2218/ijdc.v6i2.205</pub-id></mixed-citation></ref>
<ref id="B49"><label>49</label><mixed-citation publication-type="journal"><string-name><surname>Lawrence</surname>, <given-names>S</given-names></string-name>, et al. <year>2001</year>. <article-title>Persistence of web references in scientific research</article-title>. <source>Computer</source>, <volume>34</volume>: <fpage>26</fpage>&#8211;<lpage>31</lpage>. DOI: <pub-id pub-id-type="doi">10.1109/2.901164</pub-id></mixed-citation></ref>
<ref id="B50"><label>50</label><mixed-citation publication-type="journal"><string-name><surname>Ma</surname>, <given-names>X</given-names></string-name>, et al. <year>2017</year>. <article-title>Weaving a knowledge network for deep carbon science</article-title>. <source>Frontiers in Earth Science</source>, <fpage>5</fpage>. DOI: <pub-id pub-id-type="doi">10.3389/feart.2017.00036</pub-id></mixed-citation></ref>
<ref id="B51"><label>51</label><mixed-citation publication-type="journal"><string-name><surname>Mayernik</surname>, <given-names>MS</given-names></string-name>, <string-name><surname>Phillips</surname>, <given-names>J</given-names></string-name> and <string-name><surname>Nienhouse</surname>, <given-names>E</given-names></string-name>. <year>2016</year>. <article-title>Linking publications and data: Challenges, trends, and opportunities</article-title>. <source>D-Lib Magazine</source>, <fpage>22</fpage>. DOI: <pub-id pub-id-type="doi">10.1045/may2016-mayernik</pub-id></mixed-citation></ref>
<ref id="B52"><label>52</label><mixed-citation publication-type="journal"><string-name><surname>McMurry</surname>, <given-names>JA</given-names></string-name>, et al. <year>2017</year>. <article-title>Identifiers for the 21st century: How to design, provision, and reuse persistent identifiers to maximize utility and impact of life science data</article-title>. <source>PLoS Biol</source>, <volume>15</volume>: <elocation-id>e2001414</elocation-id>. DOI: <pub-id pub-id-type="doi">10.1371/journal.pbio.2001414</pub-id></mixed-citation></ref>
<ref id="B53"><label>53</label><mixed-citation publication-type="webpage"><string-name><surname>Mooney</surname>, <given-names>H</given-names></string-name> and <string-name><surname>Newton</surname>, <given-names>MP</given-names></string-name>. <year>2012</year>. <article-title>The anatomy of a data citation: Discovery, reuse, and credit</article-title>. <source>Journal of Librarianship &amp; Scholarly Communication</source>, <volume>1</volume>: <fpage>1</fpage>&#8211;<lpage>16</lpage>. <uri>https://jlsc-pub.org/articles/abstract/10.7710/2162-3309.1035/</uri> accessed 2017-10-03. DOI: <pub-id pub-id-type="doi">10.7710/2162-3309.1035</pub-id></mixed-citation></ref>
<ref id="B54"><label>54</label><mixed-citation publication-type="webpage"><collab>NISO</collab>. <year>2016</year>. <article-title>Outputs of the NISO alternative assessment metrics project: A recommended practice of the National Information Standards Organization</article-title>. NISO RP-25-2016. <uri>https://www.niso.org/publications/rp-25-2016-altmetrics</uri> accessed 2019-07-22.</mixed-citation></ref>
<ref id="B55"><label>55</label><mixed-citation publication-type="confproc"><string-name><surname>Parsons</surname>, <given-names>MA</given-names></string-name>, et al. <year>2008</year>. <article-title>Managing permafrost data: Past approaches and future directions</article-title>. <conf-name>Permafrost Ninth International Conference</conf-name> <conf-date>29 June&#8211;3 July 2008</conf-date> Proceedings, <fpage>1369</fpage>&#8211;<lpage>1374</lpage>. DOI: <pub-id pub-id-type="doi">10.5281/zenodo.3519368</pub-id></mixed-citation></ref>
<ref id="B56"><label>56</label><mixed-citation publication-type="journal"><string-name><surname>Parsons</surname>, <given-names>MA</given-names></string-name> and <string-name><surname>Fox</surname>, <given-names>PA</given-names></string-name>. <year>2013</year>. <article-title>Is data publication the right metaphor?</article-title> <source>Data Science Journal</source>, <fpage>12</fpage>. DOI: <pub-id pub-id-type="doi">10.2481/dsj.WDS-042</pub-id></mixed-citation></ref>
<ref id="B57"><label>57</label><mixed-citation publication-type="journal"><string-name><surname>Parsons</surname>, <given-names>MA</given-names></string-name> and <string-name><surname>Fox</surname>, <given-names>PA</given-names></string-name>. <year>2018</year>. <article-title>Power and persistent identifiers</article-title>. <source>International Data Week 2018</source>. DOI: <pub-id pub-id-type="doi">10.5281/zenodo.1495321</pub-id></mixed-citation></ref>
<ref id="B58"><label>58</label><mixed-citation publication-type="journal"><string-name><surname>Paskin</surname>, <given-names>N</given-names></string-name>. <year>2000</year>. <article-title>E-citations: Actionable identifiers and scholarly referencing</article-title>. <source>Learned Publishing</source>, <volume>13</volume>: <fpage>159</fpage>&#8211;<lpage>166</lpage>. DOI: <pub-id pub-id-type="doi">10.1087/09531510050145308</pub-id></mixed-citation></ref>
<ref id="B59"><label>59</label><mixed-citation publication-type="journal"><string-name><surname>Peters</surname>, <given-names>I</given-names></string-name>, et al. <year>2016</year>. <article-title>Research data explored: An extended analysis of citations and altmetrics</article-title>. <source>Scientometrics</source>, <volume>107</volume>: <fpage>723</fpage>&#8211;<lpage>744</lpage>. DOI: <pub-id pub-id-type="doi">10.1007/s11192-016-1887-4</pub-id></mixed-citation></ref>
<ref id="B60"><label>60</label><mixed-citation publication-type="book"><string-name><surname>Rauber</surname>, <given-names>A</given-names></string-name>, <string-name><surname>Asmi</surname>, <given-names>A</given-names></string-name>, <string-name><surname>van Uytvanck</surname>, <given-names>D</given-names></string-name> and <string-name><surname>Proell</surname>, <given-names>S</given-names></string-name>. <year>2015</year>. <source>Data Citation of Evolving Data: Recommendations of the Working Group on Data Citation (WGDC)</source>. <publisher-name>Research Data Alliance</publisher-name>. Accessed 2019-07-14. DOI: <pub-id pub-id-type="doi">10.15497/RDA00016</pub-id></mixed-citation></ref>
<ref id="B61"><label>61</label><mixed-citation publication-type="confproc"><string-name><surname>Schopf</surname>, <given-names>JM</given-names></string-name>. <year>2012</year>. <article-title>Treating data like software: A case for production quality data</article-title>. <conf-name>Proceedings of the Joint Conference on Digital Libraries</conf-name>, <conf-date>11&#8211;14 June 2012</conf-date>. <conf-loc>Washington DC</conf-loc>. DOI: <pub-id pub-id-type="doi">10.1145/2232817.2232846</pub-id></mixed-citation></ref>
<ref id="B62"><label>62</label><mixed-citation publication-type="journal"><string-name><surname>Silvello</surname>, <given-names>G</given-names></string-name>. <year>2018</year>. <article-title>Theory and practice of data citation</article-title>. <source>Journal of the Association for Information Science and Technology</source>, <volume>69</volume>: <fpage>6</fpage>&#8211;<lpage>20</lpage>. DOI: <pub-id pub-id-type="doi">10.1002/asi.23917</pub-id></mixed-citation></ref>
<ref id="B63"><label>63</label><mixed-citation publication-type="journal"><string-name><surname>Smith</surname>, <given-names>AM</given-names></string-name>, <string-name><surname>Katz</surname>, <given-names>DS</given-names></string-name>, <string-name><surname>Niemeyer</surname>, <given-names>KE</given-names></string-name>, <collab>FORCE11 and SCWG</collab>. <year>2016</year>. <article-title>Software citation principles</article-title>. <source>PeerJ Computer Science</source>, <volume>2</volume>: <fpage>e86</fpage>. DOI: <pub-id pub-id-type="doi">10.7717/peerj-cs.86</pub-id></mixed-citation></ref>
<ref id="B64"><label>64</label><mixed-citation publication-type="journal"><string-name><surname>Stall</surname>, <given-names>S</given-names></string-name>, et al. <year>2018</year>. <article-title>Advancing FAIR data in Earth, space, and environmental science</article-title>. <source>Eos</source>, <fpage>99</fpage>. DOI: <pub-id pub-id-type="doi">10.1029/2018EO109301</pub-id></mixed-citation></ref>
<ref id="B65"><label>65</label><mixed-citation publication-type="journal"><string-name><surname>Stockhause</surname>, <given-names>M</given-names></string-name> and <string-name><surname>Lautenschlager</surname>, <given-names>M</given-names></string-name>. <year>2017</year>. <article-title>CMIP6 data citation of evolving data</article-title>. <source>Data Science Journal</source>, <fpage>16</fpage>. DOI: <pub-id pub-id-type="doi">10.5334/dsj-2017-030</pub-id></mixed-citation></ref>
<ref id="B66"><label>66</label><mixed-citation publication-type="journal"><string-name><surname>Stuart</surname>, <given-names>D</given-names></string-name>. <year>2017</year>. <article-title>Data bibliometrics: Metrics before norms</article-title>. <source>Online Information Review</source>, <volume>41</volume>: <fpage>428</fpage>&#8211;<lpage>435</lpage>. DOI: <pub-id pub-id-type="doi">10.1108/OIR-01-2017-0008</pub-id></mixed-citation></ref>
<ref id="B67"><label>67</label><mixed-citation publication-type="journal"><collab>TGDCSP &#8211; Task Group on Data Citation Standards and Practices, CODATA-ICSTI</collab>. <year>2013</year>. <article-title>Out of cite, out of mind: The current state of practice, policy, and technology for the citation of data</article-title>. <source>Data Science Journal</source>, <volume>12</volume>: <fpage>CIDCR1</fpage>&#8211;<lpage>CIDCR75</lpage>. DOI: <pub-id pub-id-type="doi">10.2481/dsj.OSOM13-043</pub-id></mixed-citation></ref>
<ref id="B68"><label>68</label><mixed-citation publication-type="journal"><string-name><surname>Wilkinson</surname>, <given-names>MD</given-names></string-name>, et al. <year>2016</year>. <article-title>The FAIR guiding principles for scientific data management and stewardship</article-title>. <source>Scientific Data</source>, <volume>3</volume>: <elocation-id>160018</elocation-id>. DOI: <pub-id pub-id-type="doi">10.1038/sdata.2016.18</pub-id></mixed-citation></ref>
<ref id="B69"><label>69</label><mixed-citation publication-type="book"><string-name><surname>Wittenburg</surname>, <given-names>P</given-names></string-name>, <string-name><surname>Hellstr&#246;m</surname>, <given-names>M</given-names></string-name>, <string-name><surname>Zw&#246;lf</surname>, <given-names>CM</given-names></string-name>, <string-name><surname>Abroshan</surname>, <given-names>H</given-names></string-name>, et al. (eds.) <year>2017</year>. <source>Persistent identifiers: Consolidated assertions</source>. <publisher-name>Research Data Alliance</publisher-name>. DOI: <pub-id pub-id-type="doi">10.15497/RDA00027</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>