<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<!--<?xml-stylesheet type="text/xsl" href="article.xsl"?>-->
<article article-type="research-article" dtd-version="1.1" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id journal-id-type="issn">1683-1470</journal-id>
<journal-title-group>
<journal-title>Data Science Journal</journal-title>
</journal-title-group>
<issn pub-type="epub">1683-1470</issn>
<publisher>
<publisher-name>Ubiquity Press</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5334/dsj-2020-028</article-id>
<article-categories>
<subj-group>
<subject>Practice paper</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>YARD: A Tool for Curating Research Outputs</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Peer</surname>
<given-names>Limor</given-names>
</name>
<email>limor.peer@yale.edu</email>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Dull</surname>
<given-names>Joshua</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
</contrib-group>
<aff id="aff-1"><label>1</label>Institution for Social and Policy Studies, Yale University, New Haven, Connecticut, US</aff>
<aff id="aff-2"><label>2</label>Libraries, Collections, and Academic Services, The New School, New York, US</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2020-07-15">
<day>15</day>
<month>07</month>
<year>2020</year>
</pub-date>
<pub-date pub-type="collection">
<year>2020</year>
</pub-date>
<volume>19</volume>
<elocation-id>28</elocation-id>
<history>
<date date-type="received" iso-8601-date="2019-12-08">
<day>08</day>
<month>12</month>
<year>2019</year>
</date>
<date date-type="accepted" iso-8601-date="2020-06-25">
<day>25</day>
<month>06</month>
<year>2020</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright: &#x00A9; 2020 The Author(s)</copyright-statement>
<copyright-year>2020</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See <uri xlink:href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</uri>.</license-p>
</license>
</permissions>
<self-uri xlink:href="http://datascience.codata.org/articles/10.5334/dsj-2020-028/"/>
<abstract>
<p>Repositories increasingly accept research outputs and associated artifacts that underlie reported findings, leading to potential changes in the demand for data curation and repository services. This paper describes a curation tool that responds to this challenge by economizing and optimizing curation efforts. The curation tool is implemented at Yale University&#8217;s Institution for Social and Policy Studies (ISPS) as YARD. By standardizing the curation workflow, YARD helps create high quality data packages that are findable, accessible, interoperable, and reusable (FAIR) and promotes research transparency by connecting the activities of researchers, curators, and publishers through a single pipeline.</p>
</abstract>
<kwd-group>
<kwd>Data curation</kwd>
<kwd>reproducibility</kwd>
<kwd>data quality</kwd>
<kwd>workflow tool</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec>
<title>Introduction</title>
<p>YARD (Yale Application for Research Data) is an adaptable curation workflow tool that enhances research outputs and associated digital artifacts designated for archival and reuse.</p>
<sec>
<title>Quality and curation in repositories</title>
<p>The scientific principle of self-correction asks researchers to be transparent about their design, data, methods, and analysis. Transparency makes it possible for independent researchers to &#8220;reproduce reported results; test alternative specifications on the data; identify misreported or fraudulent results; reuse or adapt materials (e.g., survey instruments) for replication or extension of prior research; and better understand the interventions, measures, and context&#8221; (<xref ref-type="bibr" rid="B17">Miguel et al., 2014, p. 31</xref>). To the greatest extent possible, data and materials, including code, should be made publicly available in order to increase accountability for researcher error (<xref ref-type="bibr" rid="B6">Chambers, 2019</xref>; <xref ref-type="bibr" rid="B18">NASEM, 2018</xref>) and allow others to reproduce and confirm results (<xref ref-type="bibr" rid="B19">NASEM, 2019</xref>).</p>
<p>Universities, private and public funders, scholarly societies, journals, and other stakeholders in the scientific enterprise have been looking to data repositories to make research data, code, and other materials underlying reported findings more widely discoverable and accessible. Data repositories, however, do not apply uniform or standardized curation practices, with many offering self-deposit or opting for a minimal curation model in an attempt to appeal to busy researchers. So as the digital artifacts associated with research outputs proliferate via various repositories, they are not necessarily usable or interpretable. Stodden et al (<xref ref-type="bibr" rid="B24">2018</xref>) recently found &#8220;serious shortcomings in usability and persistence&#8221; of digital artifacts supporting publications (see also a discussion of attempts to evaluate empirical claims in published studies, <xref ref-type="bibr" rid="B16">Leek &amp; Jager, 2017</xref>). We view the ability of future users to independently understand and reuse research outputs as a key aspect of quality (see also <xref ref-type="bibr" rid="B1">Altman, 2012</xref>; <xref ref-type="bibr" rid="B2">Ashley, 2013</xref>; King, 1995; <xref ref-type="bibr" rid="B25">The Royal Society, 2012</xref>).<xref ref-type="fn" rid="n1">1</xref></p>
<p>The cost of repositories&#8217; failure to ensure the usability and interpretability, or quality, of research outputs can be great. As data science emerges as the next frontier (<xref ref-type="bibr" rid="B3">Blei &amp; Smyth, 2017</xref>; <xref ref-type="bibr" rid="B4">Burton et al., 2018</xref>), the ability to reliably use data and other digital artifacts associated with research outputs for the purpose of validating the integrity of scientific claims must be a precondition. From a practical standpoint, the community should expect that investment in research infrastructure is extended to the production of digital artifacts that can be used meaningfully. Moreover, given that, &#8220;a large number of scientific studies&#8230; suffer from the underlying computational and statistical issues&#8221; (<xref ref-type="bibr" rid="B16">Leek &amp; Jager, 2017</xref>), usable and interpretable research outputs are imperative.</p>
<p>Curation of digital research data is traditionally defined as activities that reduce threats to their long-term research value and mitigate the risk of digital obsolescence (<xref ref-type="bibr" rid="B10">DCC, n.d.</xref>). We refer to gold-standard curation as the measures taken to ensure that research outputs are independently understandable for informed reuse (<xref ref-type="bibr" rid="B21">Peer et al, 2014</xref>). Some curatorial activities &#8211; such as the periodic review of the digital integrity of a file and remedial actions to protect data from digital erosion or hardware failure &#8211; need to be ongoing. Other activities &#8211; such as code review &#8211; may be more pertinent at certain points of the data lifecycle (see <xref ref-type="bibr" rid="B15">Johnston et al., 2014</xref>, for a comprehensive list). We believe all curation activities are vital.</p>
<p>Our concern here is with the optimization of curatorial activities that enhance the quality of research outputs and associated digital materials. Ensuring the quality of these digital materials designated for long-term reuse requires effort and inevitably some cost. It has been noted that, &#8220;managing research data for quality, in one form or another, has in fact been the core responsibility of data curation since its inception as a distinct sub-discipline within the library and information sciences&#8221; (<xref ref-type="bibr" rid="B23">Sposito, 2017, p. 3</xref>). At present, however, there is no consensus in the scientific community about who is responsible for this effort: Researchers, data centers, university libraries, or data repositories? Moreover, an analysis of curation practices in general purpose and domain repositories found that review and curation of these materials are sometimes minimal or limited in scope (<xref ref-type="bibr" rid="B21">Peer et al., 2014</xref>), with self-archiving models often guaranteeing no more than bit-level preservation in order to control costs.</p>
<p>Absent a consensus on responsibilities and practices, it is not surprising that there are very few tools currently available for curators and others to manage, standardize, and share responsibility for curation activities.<xref ref-type="fn" rid="n2">2</xref> In contrast with custom tools built and used by individual repositories to accommodate their own specific curation needs and preferences, a universal tool affords other entities a way to engage with curation activities earlier in the data lifecycle. For example, laboratories can use the tool during active research, for example during data collection, to ensure that subsequent transformations to raw data are documented in sufficient detail. This can enable other researchers to trace the final analysis datasets supporting published findings back to those original raw data.</p>
<p>We describe YARD, a tool that responds to this challenge. YARD, the Yale Application for Research Data, is a workflow tool that facilitates gold-standard curation tasks. Our goal in developing YARD is to create highly curated research packages that can then be deposited into any data repository. In addition to the standard curation activities, the tool facilitates tasks for reviewing and enhancing research outputs, including verifying that data, code, and other relevant digital artifacts computationally reproduce the results those materials reportedly support. The tool also creates rich metadata about the artifacts, helping generate findable, accessible, interoperable, and reusable (FAIR)<xref ref-type="fn" rid="n3">3</xref> digital objects. By tracking curation tasks, YARD supports a transparent and documented workflow that can help researchers, curators, and publishers share responsibility for curation activities through a single pipeline. And finally, by building flexibility into the system, YARD is designed to be adapt to changing requirements and standards.</p>
</sec>
</sec>
<sec>
<title>General Description</title>
<p>The curation tool is a web-based application designed to process digital artifacts associated with research outputs, including metadata management, and to deliver highly curated research packages into designated repositories (see Figure <xref ref-type="fig" rid="F1">1</xref>). YARD is the implementation of the tool at Yale University&#8217;s Institution for Social and Policy Studies (ISPS) which began in 2018.</p>
<fig id="F1">
<label>Figure 1</label>
<caption>
<p>YARD is a workflow tool for reviewing and enhancing research outputs and delivering them into a repository.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="dsj-19-1119-g1.png"/>
</fig>
<p>Specifically, the curation tool offers two main benefits:</p>
<list list-type="order">
<list-item><p>Managing complex workflows. The workflow design helps guide depositors and curators through tasks for reviewing and enhancing research outputs. The tool tracks these curation tasks and generates rich metadata. The tool can be used to manage any updates to metadata and data, which can then be pushed out to public repositories. An advantage of the tool is that the high quality data packages it produces can be linked to different endpoints for dissemination. For archives and repositories that already do a fair amount of curation, the tool facilitates a systematic workflow with tracking and integration capabilities. For self-archiving systems that offer little or no curation, the tool can be an option for depositors, as a means of enforcing minimal documentation standards, for example. In addition, organizations with distributed expertise can use the tool to collaborate, coordinate, and standardize curation activities. For example, university library staff may be responsible for metadata generation, a statistical support unit responsible for verifying computational reproducibility of statistical analyses, and a repository responsible for assigning persistent links. An advantage of a distributed workflow model is the potential to increase the feasibility of scaling curation services without shouldering the entire cost of labor and technology.</p></list-item>
<list-item><p>Enforcing data curation standards. Specifically, the curation tool supports FAIR principles for findable, accessible, interoperable, and reusable data and other digital research artifacts. It can accommodate extending curation workflows to include additional quality checks (e.g., verification of computational reproducibility). The tool supports archival preservation policies by enforcing standards (e.g., the OAIS Reference Model requirement to clearly define roles, see <xref ref-type="bibr" rid="B5">CRL, 2007</xref>) and providing documentary evidence of such, which ISO 16363 (<xref ref-type="bibr" rid="B12">ISO, 2012</xref>) and CoreTrustSeal require (<xref ref-type="bibr" rid="B11">Dillo, 2018</xref>). As standards evolve, the tool can be configured to adapt.</p></list-item>
</list>
<sec>
<title>Technical Features</title>
<p>Key features are based on services critical to rigorous data curation: Templates for multi-file metadata creation and editing, item-level metadata creation and editing, metadata error reporting, customizable metadata exports, controlled vocabularies for selected fields, controlled vocabulary editing capabilities, record versioning, user access options, administrator and tracking controls, and a variety of content management features. The curation tool is API-enabled and modular.</p>
<p>For a minimal installation, the tool requires two open source software pieces, a web server, a database, and file storage.</p>
<p>The two open source software pieces, available on Github under a GNU Affero General Public License v3.0. (<xref ref-type="bibr" rid="B14">Iverson &amp; Smith, 2018</xref>), include the Curation Service and Curation Web application.</p>
<list list-type="simple">
<list-item><p>a) The Curation Service manages the curation workflow and logs all application events. The workflow is based on established curation steps triggered by certain user actions, and automated when possible.</p></list-item>
<list-item><p>b) The Curation Web Application provides a web interface for the Curation Service. All users &#8211; researchers depositing data and code, curators processing the outputs, and administrators &#8211; can access the curation tool through the web application.</p></list-item>
</list>
<p>The curation tool also requires,</p>
<list list-type="simple">
<list-item><p>c) A web server to host the curation tool. For the YARD implementation, the Curation Service and Web Application are installed on a Windows 2012 web server, which also hosts other software components.<xref ref-type="fn" rid="n4">4</xref></p></list-item>
<list-item><p>d) A curation database for storing the application and curation metadata.</p></list-item>
<list-item><p>e) File storage for data, codebooks, code, and all other files required for curation. The application requires storage locations for phases, a) an &#8216;original&#8217; directory for the original files and metadata comprising the research package, b) an &#8216;active&#8217; directory for copies of the files and metadata during active curation, and c) a &#8216;processed&#8217; directory for copies of the files processed and approved, as well as metadata.</p></list-item>
</list>
<p>The curation tool affords easy integration with other software or workflows. These optional components are not integral to the functioning of the curation tool but are congruent with its purpose; they can be replaced, enhanced, or left out per organization policy. Figure <xref ref-type="fig" rid="F2">2</xref> illustrates the curation tool components, their function, and relationship.</p>
<fig id="F2">
<label>Figure 2</label>
<caption>
<p>Curation tool components, required and optional.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="dsj-19-1119-g2.png"/>
</fig>
<p>At Yale, the YARD implementation of the curation tool integrates with components that provide additional desired functionality. It is configured to the requirements of ISPS and includes some proprietary software components. The YARD implementation includes,</p>
<list list-type="alpha-lower">
<list-item><p>Colectica Repository, a proprietary software developed by Colectica to create variable-level metadata extracted from SPSS, Stata, CSV, and Excel files (<xref ref-type="bibr" rid="B8">Colectica, 2016</xref>). The metadata scheme is based on the Data Documentation Initiative (DDI)<xref ref-type="fn" rid="n5">5</xref> (the tool allows adding new fields from other established metadata schemes). This software requires its own database to store the metadata.</p></list-item>
<list-item><p>ClamAV as an antivirus check for all uploaded files (<xref ref-type="bibr" rid="B26">Tiesi, 2014</xref>).</p></list-item>
<list-item><p>StatTransfer for creating plain-text copies of data files (<xref ref-type="bibr" rid="B7">Circle Systems, Inc., 2017</xref>).</p></list-item>
<list-item><p>Yale&#8217;s Persistent Linking Service to create persistent URLs (<xref ref-type="bibr" rid="B27">Yale University, 2016</xref>).</p></list-item>
</list>
<p>Table <xref ref-type="table" rid="T1">1</xref> lists the required and optional components, specifies the components used for the YARD implementation, and suggests alternative options for components where available.</p>
<table-wrap id="T1">
<label>Table 1</label>
<caption>
<p>Curation Tool Components.</p>
</caption>
<table>
<tr>
<th align="left" valign="top">Component</th>
<th align="left" valign="top">Function</th>
<th align="left" valign="top">License</th>
<th align="left" valign="top">YARD implementation</th>
<th align="left" valign="top">Alternate Component options</th>
</tr>
<tr>
<td colspan="5"><hr/></td>
</tr>
<tr>
<td colspan="5"><bold><italic>Required components</italic></bold></td>
</tr>
<tr>
<td align="left" valign="top">Curation Web Application</td>
<td align="left" valign="top">Web interface for the Curation Service</td>
<td align="left" valign="top">AGPL 3.0</td>
<td align="left" valign="top">Curation Web Application</td>
<td align="left" valign="top"></td>
</tr>
<tr>
<td align="left" valign="top">Curation Service</td>
<td align="left" valign="top">Data deposit and curation</td>
<td align="left" valign="top">AGPL 3.0</td>
<td align="left" valign="top">Curation Service</td>
<td align="left" valign="top"></td>
</tr>
<tr>
<td align="left" valign="top">Curation Database</td>
<td align="left" valign="top">Storage for curation tool data</td>
<td align="left" valign="top">Proprietary</td>
<td align="left" valign="top">Microsoft SQL</td>
<td align="left" valign="top">Postgres, MySQL</td>
</tr>
<tr>
<td align="left" valign="top">File storage</td>
<td align="left" valign="top">Storage for files</td>
<td align="left" valign="top">Yale Service</td>
<td align="left" valign="top">Network attached local service (storage@Yale)</td>
<td align="left" valign="top">Any file storage (requires read/write access)</td>
</tr>
<tr>
<td colspan="5"><bold><italic>Optional components</italic></bold></td>
</tr>
<tr>
<td align="left" valign="top">Metadata Repository</td>
<td align="left" valign="top">Captures, generates, and versions DDI Lifecycle metadata</td>
<td align="left" valign="top">Proprietary</td>
<td align="left" valign="top">Colectica Repository</td>
<td align="left" valign="top">Repository software that generates metadata</td>
</tr>
<tr>
<td align="left" valign="top">Metadata database</td>
<td align="left" valign="top">Stores variable-level metadata</td>
<td align="left" valign="top">Proprietary</td>
<td align="left" valign="top">Microsoft SQL</td>
<td align="left" valign="top">Postgres, MySQL</td>
</tr>
<tr>
<td align="left" valign="top">Anti Virus</td>
<td align="left" valign="top">Virus scan for deposited files</td>
<td align="left" valign="top">GPL</td>
<td align="left" valign="top">ClamAV</td>
<td align="left" valign="top">Any Antivirus software</td>
</tr>
<tr>
<td align="left" valign="top">File Conversion</td>
<td align="left" valign="top">Creates csv copies of data files</td>
<td align="left" valign="top">Proprietary</td>
<td align="left" valign="top">StatTransfer</td>
<td align="left" valign="top">Any statistical or custom software</td>
</tr>
<tr>
<td align="left" valign="top">Persistent Link</td>
<td align="left" valign="top">Persistent Link</td>
<td align="left" valign="top">Yale Service</td>
<td align="left" valign="top">Yale Handle service</td>
<td align="left" valign="top">Any persistent linking service</td>
</tr>
</table>
</table-wrap>
</sec>
<sec>
<title>User Roles</title>
<p>All users are required to create an account in the curation tool.<xref ref-type="fn" rid="n6">6</xref> Users are assigned one of the following roles: Depositor, Curator, or Organization Admin. Users of the curation tool will experience a different workflow and have access to different features depending on their role in the system. Below, we discuss the curation workflow through the lens of the three main roles.</p>
<sec>
<title>Depositor</title>
<p>By default, all users are Depositors. A Depositor can submit data, code, and other research outputs that comprise a Catalog Record. A Depositor can be a researcher affiliated with an organization hosting the curation tool or one of the organization&#8217;s staff members. A Depositor creating a new record will add study-level metadata (e.g., author, title, sample size, field dates, etc.), upload all related files, add file-level metadata, and finally submit for curation. Figure <xref ref-type="fig" rid="F3">3</xref> is an example of uploaded files associated with a sample study.</p>
<fig id="F3">
<label>Figure 3</label>
<caption>
<p>The Depositor view of the file list after initial upload, as seen in the user interface.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="dsj-19-1119-g3.png"/>
</fig>
<p>Each file is checked by Clam AntiVirus, assigned a universally unique identifier (UUID), and deposited into the &#8216;original&#8217; directory, where copies of all the raw data are kept. A Depositor can update the record if there are changes or new additions to a study and re-submit for curation. Each version submitted is preserved in the &#8216;processed&#8217; directory so no data are lost.</p>
</sec>
<sec>
<title>Organization Admin and Curator</title>
<p>Once a Depositor submits a study for curation, copies are made and stored in the &#8216;active&#8217; directory and a notification is sent to the Organization Admin. As noted, the Organization Admin has permissions to edit the organization&#8217;s settings, including setting the domain name, assigning storage locations, specifying a repository destination, adding a deposit agreement, managing roles and permissions, and other technical settings. The Admin assigns a Curator to each Catalog Record and approves publication once curation is complete.</p>
<p>When notified of a new record, the Curator will complete all curation tasks which the tool automatically assigns based on the file types (for example, a data file will have different curation steps than a codebook). Figure <xref ref-type="fig" rid="F4">4</xref> is an example of the Curator&#8217;s review panel which lists curation tasks. Once curation is complete, the Curator will submit the Catalog Record for publication approval by the Admin.</p>
<fig id="F4">
<label>Figure 4</label>
<caption>
<p>The Curator view of the curation tasks for the assigned record, as viewed in the user interface.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="dsj-19-1119-g4.png"/>
</fig>
<p>A record publication approval triggers a series of events on the server side. A plain-text, preservation copy (e.g., .csv) is created of certain proprietary data files (e.g., Stata .dta) and added to the &#8216;processed&#8217; directory. Any updates or additions to the study-level metadata are synced in the Curation Database. Changes to the files are versioned by Git, which is built into the Curation Service. The Curation Database also stores the details about each completed curation step including the date, time, and which user completed the step. The tool provides a full history log for each Catalog Record. If configured, updates to the data files and variables are synced to the Metadata Database. The optional Colectica Repository software uses metadata from both the Curation &amp; Metadata Databases to create a detailed metadata file using the Data Documentation Initiative (DDI) 3.2 schema. Finally, the Catalog Record, and each file marked as &#8216;public&#8217;, are assigned a persistent link. YARD uses Yale&#8217;s in-house handle service to generate these links, but integration with other services is possible.</p>
</sec>
</sec>
<sec>
<title>The Curation Workflow</title>
<p>The curation workflow as implemented in YARD is designed to the specifications of ISPS at Yale University. The workflow is based on the Inter-university Consortium for Political and Social Research (<xref ref-type="bibr" rid="B13">ICPSR, n.d.</xref>) pipeline and adapted for quantitative research output from randomized controlled trials (RCTs) in the social sciences (<xref ref-type="bibr" rid="B21">Peer et al, 2014</xref>). For example, YARD prompts Curators to review whether documentation and contextual information necessary for long-term usability (e.g., a codebook, a readme file) are included. The curation workflow has been further enhanced to include tasks for reviewing code and statistical analysis to obtain verification of computational reproducibility (<xref ref-type="bibr" rid="B22">Peer &amp; Wykstra, 2015</xref>; <xref ref-type="bibr" rid="B20">Peer, 2017</xref>). For example, YARD guides Curators to review code files &#8211; statistical and other programming scripts &#8211; by verifying that the code executes and that the published scientific results can be computationally reproduced with the given code and data. The workflow was developed with input from potential users at the Odum Institute Data Archive at the University of North Carolina, Chapel Hill and the Cornell Institute for Social and Economic Research at Cornell University.</p>
</sec>
<sec>
<title>Integration with a Repository</title>
<p>The curation tool is not a replacement for a repository in so far as it is not meant to be an access point for other scholars or the general public. End users can only access records and files processed through the curation tool if they are ingested into another system such as a data repository or archive. The YARD implementation is currently designed to integrate with Drupal and provides access to processed records via the ISPS Data Archive.<xref ref-type="fn" rid="n7">7</xref> Organizations can determine a preferred means of dissemination based on their own infrastructure.</p>
<p>There are two methods for integrating processed records into other software or workflows: the &#8216;processed&#8217; directory and the Extensible Markup Language (XML) feed. The &#8216;processed&#8217; directory is a compressed directory created by the curation tool when each catalog record is finalized and approved. The zipped archive, as seen expanded in Figure <xref ref-type="fig" rid="F5">5</xref>, is structured to match BagIt specifications.<xref ref-type="fn" rid="n8">8</xref> It contains a file manifest with MD5 checksums, a handle map, and a &#8216;data&#8217; directory containing the curated files and the application-generated DDI file. This DDI file contains the metadata necessary to ingest studies into another system or repository. Since studies can be re-curated and re-approved, a unique archive directory is created for each instance of publication (e.g., a study reviewed and finalized a second time will have two archive directories, one for each version).</p>
<fig id="F5">
<label>Figure 5</label>
<caption>
<p>An example archive directory for a catalog record after curation and final approval.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="dsj-19-1119-g5.png"/>
</fig>
<p>The second option for disseminating records is the XML feed, which is created when a record is approved for publication. The feed contains metadata about each study and any associated files processed through the curation tool. Figure <xref ref-type="fig" rid="F6">6</xref> shows a sample of the XML feed. Only studies and files marked as &#8216;public&#8217; will appear in the feed. The feed includes the persistent link for each file, so files can be downloaded or ingested via the feed. The feed can be ingested into any system with a XML mapping option. In the YARD implementation, the XML feed is ingested into a Drupal site using the Drupal feed importers module,<xref ref-type="fn" rid="n9">9</xref> whereby each XML element is mapped to a specific Drupal field.</p>
<fig id="F6">
<label>Figure 6</label>
<caption>
<p>An example XML feed produced by YARD showing the metadata for a catalog record.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="dsj-19-1119-g6.png"/>
</fig>
</sec>
<sec>
<title>Customizability</title>
<p>The curation tool is designed to be fully customizable. That includes configuring the curation workflow such that curation tasks can be adjusted. For example, other curation frameworks could be applied (e.g., the Data Curation Network&#8217;s CURATED checklist, see <xref ref-type="bibr" rid="B9">DCN, 2018</xref>) or specific tasks can be made optional or dropped altogether (e.g., checking for the presence of personally-identifiable information). These and other changes to the application code&#8211;changing user roles (e.g., the Depositor role may be eliminated if a repository only allows Curators to deposit, or other roles can be added), changing the metadata schema as appropriate to other disciplines, customizing study-level information (e.g., adding a geolocation field), and more&#8211;can be done by a skilled developer.</p>
<p>Other customization involves changes to admin settings in the application or editing config files on the web server. For example, within the application, administrators can customize file storage locations, edit or turn off persistent link minting, and integrate with other software to automate tasks where possible (see Table <xref ref-type="table" rid="T1">1</xref>). Admin settings also allow customization of email notifications and permissions to create new users accounts. Finally, admin access to the web server grants permissions to edit config files in order to customize functions such as changing database paths, limit or increase allowable upload file size, and customize error reporting.</p>
<p>Given the diversity of research practices and products, the tool is designed to be modular and built with interchangeable components, such that it could be customized by an organization to meet its specifications, requirements, and policies.</p>
</sec>
<sec>
<title>Availability</title>
<p>YARD is currently supported at Yale with local IT resources and infrastructure, including admins who deploy the full stack and monitor and maintain the web and database servers. Any organization that assumes responsibility for the curation of research outputs&#8211;for example, a repository, a research lab, an academic research center, or a library&#8211;can have a local installation by compiling from the source code. Access to the code and comprehensive documentation are available in a public repository (<xref ref-type="bibr" rid="B14">Iverson &amp; Smith, 2018</xref>). As an open source project, it is our hope that interested parties will join us in supporting and improving the software in accordance with best practice governance models in the academic Open Source community.</p>
</sec>
</sec>
<sec>
<title>Discussion and Conclusions</title>
<p>This paper describes a policy-driven adaptable workflow tool supporting the archival and dissemination of high-quality research outputs. The tool is designed to increase the potential for long-term usability by creating high quality and FAIR-compliant data packages. The essential design principles applied to this tool is modularity and open source. The tool also promotes research transparency by connecting the activities of researchers, curators, and publishers through a single pipeline. Our vision is for this tool to be used by organizations committed to both rigorous research practices and high-quality output. We believe this project is a significant step toward the &#8220;development of more generic tools and processes for validating and improving various aspects of data quality,&#8221; as called by Digital Curation Centre Director, Kevin Ashley (<xref ref-type="bibr" rid="B2">2013</xref>).</p>
<p>YARD addresses variability in research output quality by helping economize and standardize curation efforts and services. It achieves that by,</p>
<list list-type="order">
<list-item><p>Providing a workflow in which curation activities can be managed, tracked, inspected, standardized, and shared and,</p></list-item>
<list-item><p>Enabling implementation of quality standards and policies aligned with making research outputs more usable and interpretable in the long term and deploying a design approach that facilitates accommodating new conditions and integrating with improved tools.</p></list-item>
</list>
<p>Developing YARD was the collaborative effort of several groups at Yale and Colectica. The team made use of project management tools to communicate with the developers, track software bugs, and document the software development process. At Yale, good working relationships with partners in Yale Information Technology Services and Yale University Library IT were essential to the project&#8217;s success in all steps of development. Looking back at the trajectory of the project&#8217;s development, we recognize that, as with many software development projects, we were subject to tightly resourced environments that presented a challenge to well-intentioned but sometime compromised efforts to test and deploy the tool within scheduled timelines and to assume local project ownership beyond the initial Colectica development. A more agile approach to deployment and testing of the software could have mitigated the consequences of some legacy decisions made at the project&#8217;s inception, such as, hosting the software on Yale ITS managed infrastructure (which provided automated server backups and security management but required additional coordination across departments) as opposed to a cloud service like AWS (which would give us more flexibility and control but require additional internal resources). Despite a lack of funding beyond the initial development and unforeseen delays, we have confidence in YARD&#8217;s sound fundamentals and potential to contribute to standardized, efficient, and transparent curation.</p>
<p>Future improvements to the software may include developing an API to allow further integration of published records with various workspaces or repository destinations. The curation log generated by the tool may be mined for information about curation tasks to inform staffing needs and educational efforts relating to research data management and curation. The curation tool&#8217;s version tracking and UUID capabilities may be used to track the evolution of digital objects throughout the research lifecycle, from creation to publication or archival, and to link them to other systems, such as institutional sponsored projects record keeping. Related, other methods of authentication may be implemented to allow seamless integration with other systems. We urge the community to take advantage of the open source software. For now, we are confident that the curation tool provides a framework and a method for enhancing the digital artifacts underpinning scientific research &#8211; something that research institutions, repositories and archives, and publishers have a vested interest in.</p>
</sec>
</body>
<back>
<fn-group>
<fn id="n1"><p>Other aspects of quality, such as the accuracy and validity of the data or the soundness of analytic choices captured in code, are outside our definition of quality. We defer these evaluations to the scientific community.</p></fn>
<fn id="n2"><p>We note the Dataverse Data Curation Tool, currently in development, which is meant to be used with Dataverse. See: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/scholarsportal/Dataverse-Data-Curation-Tool">https://github.com/scholarsportal/Dataverse-Data-Curation-Tool</ext-link> (accessed 2020 January 2).</p></fn>
<fn id="n3"><p>See: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.go-fair.org/fair-principles/">https://www.go-fair.org/fair-principles/</ext-link> (accessed 2019 November 27).</p></fn>
<fn id="n4"><p>The components are written in C#. The current version of the curation tool runs on a Windows server and the Web framework is .aspnet.</p></fn>
<fn id="n5"><p>See: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://ddialliance.org/">https://ddialliance.org/</ext-link> (accessed 2019 December 30).</p></fn>
<fn id="n6"><p>The YARD implementation uses a password log in method; future development can include other methods of authentication.</p></fn>
<fn id="n7"><p>See: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://isps.yale.edu/research/data">https://isps.yale.edu/research/data</ext-link> (accessed 2019 November 18).</p></fn>
<fn id="n8"><p>See: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://www.dcc.ac.uk/resources/external/bagit-library">http://www.dcc.ac.uk/resources/external/bagit-library</ext-link> (accessed 2019 November 27).</p></fn>
<fn id="n9"><p>See: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.drupal.org/project/feed_import">https://www.drupal.org/project/feed_import</ext-link> (accessed 2019 November 18).</p></fn>
</fn-group>
<ack>
<title>Acknowledgements</title>
<p>We are grateful to our reviewers for helpful comments about this manuscript. We thank Yale University Library and Yale Information Technology Services for supporting this project. We give special thanks to Ann Green for contributing to the conceptualization of this project, to the team at Innovations for Poverty Action for early planning, to Themba Flowers for contributing to the deployment at Yale, and to Mike Friscia, Eric James, and Robert Wolfe for critical support at Yale. We also thank Jeremy Iverson and Dan Smith at Colectica for their dedicated development work, and Thu-Mai Christian at the Odum Institute Data Archive and Florio Arguillas at the Cornell Institute for Social and Economic Research for participating in early testing. This project would not have been possible without the support of the Institution for Social and Policy Studies at Yale University and its commitment to upholding and implementing the highest standards in academic research. L.P. acknowledges funding from Innovations for Poverty Action.</p>
</ack>
<sec>
<title>Competing Interests</title>
<p>The authors have no competing interests to declare.</p>
</sec>
<ref-list>
<ref id="B1"><label>1</label><mixed-citation publication-type="webpage"><string-name><surname>Altman</surname>, <given-names>M</given-names></string-name>. <year>2012</year>. <chapter-title>Mitigating Threats to Data Quality Throughout the Curation Lifecycle</chapter-title>. In: <string-name><surname>Marchionini</surname>, <given-names>G</given-names></string-name>, <string-name><surname>Lee</surname>, <given-names>CA</given-names></string-name>, <string-name><surname>Bowden</surname>, <given-names>H</given-names></string-name> and <string-name><surname>Lesk</surname>, <given-names>M</given-names></string-name> (eds.), <source>Curating for Quality: Ensuring Data Quality to Enable New Science</source>. Final Report: <publisher-name>Invitational Workshop Sponsored by the National Science Foundation</publisher-name>, <month>Sept</month>. <day>10&#8211;11</day>, 2012, <publisher-loc>Arlington, VA</publisher-loc>. <uri>https://ils.unc.edu/callee/curating-for-quality.pdf</uri>.</mixed-citation></ref>
<ref id="B2"><label>2</label><mixed-citation publication-type="journal"><string-name><surname>Ashley</surname>, <given-names>K</given-names></string-name>. <year>2013</year>. <article-title>Data Quality and Curation</article-title>. <source>Data Science Journal</source>, <volume>12</volume>: <fpage>GRDI65</fpage>&#8211;<lpage>GRDI68</lpage>. DOI: <pub-id pub-id-type="doi">10.2481/dsj.GRDI-011</pub-id></mixed-citation></ref>
<ref id="B3"><label>3</label><mixed-citation publication-type="journal"><string-name><surname>Blei</surname>, <given-names>DM</given-names></string-name> and <string-name><surname>Smyth</surname>, <given-names>P</given-names></string-name>. <year>2017</year>. <article-title>Science and Data Science</article-title>. <source>Proceedings of the National Academy of Sciences</source>, <volume>114</volume>(<issue>33</issue>): <fpage>8689</fpage>&#8211;<lpage>8692</lpage>. DOI: <pub-id pub-id-type="doi">10.1073/pnas.1702076114</pub-id></mixed-citation></ref>
<ref id="B4"><label>4</label><mixed-citation publication-type="webpage"><string-name><surname>Burton</surname>, <given-names>M</given-names></string-name>, <string-name><surname>Lyon</surname>, <given-names>L</given-names></string-name>, <string-name><surname>Erdmann</surname>, <given-names>C</given-names></string-name> and <string-name><surname>Tijerina</surname>, <given-names>B</given-names></string-name>. <year>2018</year>. <source>Shifting to Data Savvy: The Future of Data Science in Libraries</source>. Project Report. <publisher-name>University of Pittsburgh</publisher-name>, <publisher-loc>Pittsburgh, PA</publisher-loc>. <uri>http://d-scholarship.pitt.edu/id/eprint/33891</uri>.</mixed-citation></ref>
<ref id="B5"><label>5</label><mixed-citation publication-type="webpage"><collab>Center for Research Libraries (CRL)</collab>. <year>2007</year>. <source>Trustworthy Repositories Audit &amp; Certification: Criteria and Checklist</source>. <uri>http://www.crl.edu/sites/default/files/d6/attachments/pages/trac_0.pdf</uri> (accessed on 2020 April 2).</mixed-citation></ref>
<ref id="B6"><label>6</label><mixed-citation publication-type="journal"><string-name><surname>Chambers</surname>, <given-names>K</given-names></string-name>, et al. <year>2019</year>. <article-title>Towards Minimum Reporting Standards for Life Scientists</article-title>. <source>MetaArXiv</source>. <month>April</month> <day>30</day>. DOI: <pub-id pub-id-type="doi">10.31222/osf.io/9sm4x</pub-id></mixed-citation></ref>
<ref id="B7"><label>7</label><mixed-citation publication-type="webpage"><collab>Circle Systems, Inc</collab>. <year>2017</year>. <article-title>Stat/Transfer (Version 14). [computer software]</article-title>. Available from <uri>https://stattransfer.com/</uri>.</mixed-citation></ref>
<ref id="B8"><label>8</label><mixed-citation publication-type="webpage"><collab>Colectica</collab>. <year>2016</year>. <article-title>Colectica Repository (Version 5.0.4236). [computer software]</article-title>. Available from <uri>https://www.colectica.com/software/repository/</uri>.</mixed-citation></ref>
<ref id="B9"><label>9</label><mixed-citation publication-type="webpage"><collab>Data Curation Network</collab>. <year>2018</year>. <source>Checklist of CURATED Steps Performed by the Data Curation Network</source>. <uri>http://z.umn.edu/curate</uri>. (also found at <uri>https://datacurationnetwork.org</uri>)</mixed-citation></ref>
<ref id="B10"><label>10</label><mixed-citation publication-type="webpage"><collab>Digital Curation Centre (DCC)</collab>. n.d. <source>What is Digital Curation?</source> <uri>http://www.dcc.ac.uk/digital-curation/what-digital-curation</uri> (accessed on 2019 December 30).</mixed-citation></ref>
<ref id="B11"><label>11</label><mixed-citation publication-type="journal"><string-name><surname>Dillo</surname>, <given-names>I</given-names></string-name> and <string-name><surname>Leeuw</surname>, <given-names>L</given-names></string-name>. <year>2018</year>. <article-title>CoreTrustSeal</article-title>. <source>Mitteilungen der Vereinigung &#214;sterreichischer Bibliothekarinnen &amp; Bibliothekare</source>, <volume>71</volume>(<issue>1</issue>): <fpage>162</fpage>&#8211;<lpage>170</lpage>. DOI: <pub-id pub-id-type="doi">10.31263/voebm.v71i1.1981</pub-id></mixed-citation></ref>
<ref id="B12"><label>12</label><mixed-citation publication-type="webpage"><collab>International Organization for Standardization (ISO)</collab>. <year>2012</year>. <source>ISO 16363:2012 &#8211; Space Data and Information Transfer Systems &#8211; Audit and Certification of Trustworthy Digital Repositories</source>. <publisher-loc>Geneva</publisher-loc>: <publisher-name>International Organization for Standardization</publisher-name>. <uri>http://www.iso.org/iso/iso_catalogue/catalogue_tc/catalogue_detail.htm?csnumber=56510</uri> (accessed on 2020 April 2).</mixed-citation></ref>
<ref id="B13"><label>13</label><mixed-citation publication-type="webpage"><collab>Inter-university Consortium for Political and Social Research (ICPSR)</collab>. n.d. <source>Data Enhancement</source>. <uri>https://www.icpsr.umich.edu/icpsrweb/content/datamanagement/lifecycle/ingest/enhance.html</uri> (accessed on 2020 April 2).</mixed-citation></ref>
<ref id="B14"><label>14</label><mixed-citation publication-type="webpage"><string-name><surname>Iverson</surname>, <given-names>J</given-names></string-name> and <string-name><surname>Smith</surname>, <given-names>D</given-names></string-name>. <year>2020</year>, <month>January</month> <day>8</day>. <article-title>Colectica/Curation: Initial Release (Version v0.9). Zenodo</article-title>. DOI: <pub-id pub-id-type="doi">10.5281/zenodo.3600615</pub-id></mixed-citation></ref>
<ref id="B15"><label>15</label><mixed-citation publication-type="journal"><string-name><surname>Johnston</surname>, <given-names>LR</given-names></string-name>, et al. <year>2018</year>. <article-title>How Important Are Data Curation Activities to Researchers? Gaps and Opportunities for Academic Libraries</article-title>. <source>Journal of Librarianship and Scholarly Communication</source>, <volume>6</volume>(<issue>General Issue</issue>): <fpage>eP2198</fpage>. DOI: <pub-id pub-id-type="doi">10.7710/2162-3309.2198</pub-id></mixed-citation></ref>
<ref id="B16"><label>16</label><mixed-citation publication-type="journal"><string-name><surname>Leek</surname>, <given-names>JT</given-names></string-name> and <string-name><surname>Jager</surname>, <given-names>LR</given-names></string-name>. <year>2017</year>. <article-title>Is Most Published Research Really False?</article-title> <source>Annual Review of Statistics and Its Application</source>, <volume>4</volume>(<issue>1</issue>): <fpage>109</fpage>&#8211;<lpage>122</lpage>. DOI: <pub-id pub-id-type="doi">10.1146/annurev-statistics-060116-054104</pub-id></mixed-citation></ref>
<ref id="B17"><label>17</label><mixed-citation publication-type="webpage"><string-name><surname>Miguel</surname>, <given-names>E</given-names></string-name>, et al. <year>2014</year>. <article-title>Promoting transparency in social science research</article-title>. <source>Science</source>, <volume>343</volume>(<issue>6166</issue>): <fpage>30</fpage>&#8211;<lpage>31</lpage>. <uri>http://science.sciencemag.org/content/343/6166/30</uri>. DOI: <pub-id pub-id-type="doi">10.1126/science.1245317</pub-id></mixed-citation></ref>
<ref id="B18"><label>18</label><mixed-citation publication-type="book"><collab>National Academies of Sciences, Engineering, and Medicine</collab>. <year>2018</year>. <source>Open Science by Design: Realizing a Vision for 21st Century Research</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>The National Academies Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.17226/25116</pub-id></mixed-citation></ref>
<ref id="B19"><label>19</label><mixed-citation publication-type="book"><collab>National Academies of Sciences, Engineering, and Medicine</collab>. <year>2019</year>. <source>Reproducibility and Replicability in Science</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>The National Academies Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.17226/25303</pub-id></mixed-citation></ref>
<ref id="B20"><label>20</label><mixed-citation publication-type="webpage"><string-name><surname>Peer</surname>, <given-names>L</given-names></string-name>. <year>2017</year>. <chapter-title>Enabling Scientific Reproducibility with Data Curation and Code Review</chapter-title>. In: <string-name><surname>Johnston</surname>, <given-names>L</given-names></string-name> (ed.), <source>Curating Research Data Volume Two: A Handbook of Current Practice</source>. <publisher-loc>Chicago, Illinois</publisher-loc>: <publisher-name>Association of College and Research Libraries</publisher-name>. <uri>http://hdl.handle.net/11299/185335</uri>.</mixed-citation></ref>
<ref id="B21"><label>21</label><mixed-citation publication-type="journal"><string-name><surname>Peer</surname>, <given-names>L</given-names></string-name>, <string-name><surname>Green</surname>, <given-names>A</given-names></string-name> and <string-name><surname>Stephenson</surname>, <given-names>E</given-names></string-name>. <year>2014</year>. <article-title>Committing to Data Quality Review</article-title>. <source>International Journal of Digital Curation</source>, <volume>9</volume>(<issue>1</issue>): <fpage>263</fpage>&#8211;<lpage>291</lpage>. DOI: <pub-id pub-id-type="doi">10.2218/ijdc.v9i1.317</pub-id></mixed-citation></ref>
<ref id="B22"><label>22</label><mixed-citation publication-type="webpage"><string-name><surname>Peer</surname>, <given-names>L</given-names></string-name> and <string-name><surname>Wykstra</surname>, <given-names>S</given-names></string-name>. <year>2015</year>. <article-title>New Curation Software: Step-by-Step Preparation of Social Science Data and Code for Publication and Preservation</article-title>. <source>IASSIST Quarterly</source>, <volume>39</volume>(<issue>4</issue>): <fpage>6</fpage>&#8211;<lpage>13</lpage>. <uri>https://iassistquarterly.com/index.php/iassist/article/view/902</uri> (accessed on 2019 October 31). DOI: <pub-id pub-id-type="doi">10.29173/iq902</pub-id></mixed-citation></ref>
<ref id="B23"><label>23</label><mixed-citation publication-type="confproc"><string-name><surname>Sposito</surname>, <given-names>FA</given-names></string-name>. <year>2017</year>. <article-title>What Do Data Curators Care About? Data Quality, User Trust, and the Data Reuse Plan</article-title>. <conf-name>Paper presented at: IFLA 2017 Meeting</conf-name>, <conf-loc>Wroc&#322;aw, Poland</conf-loc>. <uri>http://library.ifla.org/id/eprint/1797</uri>.</mixed-citation></ref>
<ref id="B24"><label>24</label><mixed-citation publication-type="journal"><string-name><surname>Stodden</surname>, <given-names>V</given-names></string-name>, <string-name><surname>Seiler</surname>, <given-names>J</given-names></string-name> and <string-name><surname>Ma</surname>, <given-names>Z</given-names></string-name>. <year>2018</year>. <article-title>An Empirical Analysis of Journal Policy Effectiveness for Computational Reproducibility</article-title>. <source>Proceedings of the National Academy of Sciences</source>, <volume>115</volume>(<issue>11</issue>): <fpage>2584</fpage>&#8211;<lpage>2589</lpage>. DOI: <pub-id pub-id-type="doi">10.1073/pnas.1708290115</pub-id></mixed-citation></ref>
<ref id="B25"><label>25</label><mixed-citation publication-type="webpage"><collab>The Royal Society</collab>. <year>2012</year>. <source>Science as Open Enterprise</source>. <uri>https://royalsociety.org/~/media/policy/projects/sape/2012-06-20-saoe.pdf</uri> (accessed on 2019 December 4).</mixed-citation></ref>
<ref id="B26"><label>26</label><mixed-citation publication-type="webpage"><string-name><surname>Tiesi</surname>, <given-names>G</given-names></string-name>. <year>2014</year>. <article-title>ClamAV Native Win32 Port (Version 0.98.4). [computer software]</article-title>. Available from <uri>https://github.com/clamwin/clamav-win32-old/tree/clamav-0.98</uri>.</mixed-citation></ref>
<ref id="B27"><label>27</label><mixed-citation publication-type="webpage"><collab>Yale University</collab>. <year>2016</year>. <article-title>Yale Persistent Linking Service &#8211; Web Service (Version 1.0). [computer software]</article-title>. Available from <uri>http://link.its.yale.edu/ypls-ws/PersistentLinking?wsdl</uri>.</mixed-citation></ref>
</ref-list>
</back>
</article>