23 papers · ranked by Valyu relevance
Heather A. Piwowar, Todd J. Vision, Xiaolei Huang
Background. Attribution to the original contributor upon reuse of published data is important both as a reward for data creators and to document the provenance of research findings. Previous studies have found that papers with publicly available datasets receive a higher number of citations than similar studies without…
Alisa Bokulich, Wendy Parker
We critically engage two traditional views of scientific data and outline a novel philosophical view that we call the pragmatic-representational (PR) view of data. On the PR view, data are representations that are the product of a process of inquiry, and they should be evaluated in terms of their adequacy or fitness…
Christopher J. Markiewicz, Krzysztof J. Gorgolewski, Franklin Feingold, Ross Blair + 8 more
The sharing of research data is essential to ensure reproducibility and maximize the impact of public investments in scientific research. Here we describe OpenNeuro, a BRAIN Initiative data archive that provides the ability to openly share data from a broad range of brain imaging data types following the FAIR…
Renata Gonçalves Curty, Kevin Crowston, Alison Specht, Bruce W. Grant + 2 more
'Bruce W. Grant' 'Elizabeth D. Dalton' 'Cassidy Rose Sugimoto'] The value of sharing scientific research data is widely appreciated, but factors that hinder or prompt the reuse of data remain poorly understood. Using the Theory of Reasoned Action, we test the relationship between the beliefs and attitudes of scientists…
Laura Koesten, Pavlos Vougiouklis, Elena Simperl, Paul Groth
Title: Summary The web provides access to millions of datasets that can have additional impact when used beyond their original context. We have little empirical insight into what makes a dataset more reusable than others and which of the existing guidelines and frameworks, if any, make a difference. In this paper, we…
Marcel LaFlamme, Marion Poetz, Daniel Spichtinger, Sergi Fàbregues
Considerable resources are being invested in strategies to facilitate the sharing of data across domains, with the aim of addressing inefficiencies and biases in scientific research and unlocking potential for science-based innovation. Still, we know too little about what determines whether scientific researchers…
Seth Carbon, Robin Champieux, Julie McMurry, Lilly Winfree + 2 more
Data is the foundation of science, and there is an increasing focus on how data can be reused and enhanced to drive scientific discoveries. However, most seemingly “open data” do not provide legal permissions for reuse and redistribution. Not being able to integrate and redistribute our collective data resources blocks…
Bradly Alicea
Participation in Open Data initiatives require two semi-independent actions: the sharing of data produced by a researcher or group, and a consumer of shared data. Consumers of shared data range from people interested in validating the results of a given study to transformers of the data. These transformers can add…
Serena Bonaretti, Egon Willighagen
Data sharing and reuse are crucial to enhance scientific progress and maximize return of investments in science. Although attitudes are increasingly favorable, data reuse remains difficult for lack of infrastructures, standards, and policies. The FAIR (findable, accessible, interoperable, reusable) principles aim to…
George Alter, Alejandra Gonzalez-Beltran, Lucila Ohno-Machado, Philippe Rocca-Serra
This article presents elements in the Data Tags Suite (DATS) metadata schema describing data access, data use conditions, and consent information. DATS is a product of the bioCADDIE Project, which created a data discovery index for searching across all types of biomedical data. The “access and use” metadata items in…
Authors not listed
The discoverability and reusability of data is critical for machine learning to drive new discovery in the chemical sciences, and the ‘FAIR Guiding Principles for scientific data management and stewardship’ provide a measurable set of guidelines that can be used to ensure the accessibility of reusable data. We…
David J. Hand
Ready data availability, cheap storage capacity, and powerful tools for extracting information from data have the potential to significantly enhance the human condition. However, as with all advanced technologies, this comes with the potential for misuse. Ethical oversight and constraints are needed to ensure that an…
Peter Müllner, Elisabeth Lex, Markus Schedl, Dominik Kowald
User-based KNN recommender systems (UserKNN) utilize the rating data of a target user's nearest neighbors in the recommendation process. This, however, increases the privacy risk of the neighbors, since the recommendations could expose the neighbors' rating data to other users or malicious parties. To reduce this risk…
Authors not listed
Computational explorations of reaction mechanism which support and guide experimental efforts has become a key tool in the organic and inorganic chemistry community. This Perspective addresses key challenges and best practices for generating reliable, reproducible, and reusable data for quantum chemical calculations of…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…
Arpit Narechania, Fan Du, Atanu R. Sinha, Ryan A. Rossi + 5 more
'Jane Hoffswell' 'Shunan Guo' 'Eunyee Koh' 'Shamkant B. Navathe' 'Alex Endert'] Selecting relevant data subsets from large, unfamiliar datasets can be difficult. We address this challenge by modeling and visualizing two kinds of auxiliary information: (1) quality – the validity and appropriateness of data required to…
Yannick Ureel, Maarten R. Dobbelaere, Yi Ouyang, Kevin De Ras + 3 more
By combining machine learning with design of experiments, so-called active machine learning, more efficient and cheaper research can be conducted. Machine learning algorithms are more flexible, and are better at investigating the processes spanning all length scales of chemical engineering. While the active machine…
Yu-Chieh Huang, Pierre Tremouilhac, Stefan Kuhn, Pei-Chi Huang + 6 more
A method for data review in chemical sciences with a focus on data for the characterization of synthetic molecules is described. As current procedures for data curation in chemistry rely almost exclusively on manual checking or peer reviewing, a (semi-)automatic procedure for the evaluation of data assigned to…
Li, Zhi, Zhang, Lei + 10 more
The data circulation is a complex scenario involving a large number of participants and different types of requirements, which not only has to comply with the laws and regulations, but also faces multiple challenges in technical and business areas. In order to systematically and comprehensively address these issues, it…
Michaela Regneri, Julia S. Georgi, Jurij Kost, Niklas Pietsch + 1 more
'Sabine Stamm'] > Abstract. We present an approach to compute the monetary value of individual data points, in context of an automated decision system. The proposed method enables us to explore and implement a paradigm of data minimalism for large-scale machine learning systems. Data minimalistic implementations…
Kevin Maik Jablonka, Andrew S. Rosen, Aditi S. Krishnapriyan, Berend Smit
The space of all plausible materials for a given application is so large that it cannot be explored using a brute-force approach. This is, in particular, the case for reticular chemistry which provides materials designers with a practically infinite playground on different length scales. One promising approach to guide…
Daochen Zha, Zaid Pervaiz Bhat, Kwei-Herng Lai, Fan Yang + 3 more
'Zhimeng Jiang' 'Shaochen Zhong' 'Xia Hu'] Artificial Intelligence (AI) is making a profound impact in almost every domain. A vital enabler of its great success is the availability of abundant and high-quality data for building machine learning models. Recently, the role of data in AI has been significantly magnified…
Authors not listed
The Digital Catalysis Platform (DigCat) is a pioneering integration of big data and AI tailored for catalysis materials research. It encompasses over 400,000 experimental performance data for electro-, thermo-, and photocatalysts, alongside more than 300,000 catalyst structures. DigCat provides dynamic data…