27 papers · ranked by Valyu relevance
Leigh Dodds
The FAIR principles need to be applied in context. To do that, we need to understand both the needs of data users and the characteristics of the data to be shared. This Opinion introduces ten different dataset archetypes that can be used to inform plans for how data are to be accessed, used, and shared.
Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis + 3 more
'Luis Ibáñez' 'Emilia Kacprzak' 'Paul Groth'] Abstract Generating value from data requires the ability to find, access and make sense of datasets. There are many efforts underway to encourage data sharing and reuse, from scientific publishers asking authors to submit data alongside manuscripts to data marketplaces…
Lucina Hackman, Pauline Mack, Hervé Ménard
Data underpinning science have become one of the most precious assets in research, and while the principles of FAIR (Findable, Accessible, Interoperable and Reusable) have been put forward as a guide to how to approach data handling, data sharing and long-term storage still remain a challenge for many research areas…
Susanna-Assunta Sansone, Alejandra Gonzalez-Beltran, Philippe Rocca-Serra, George Alter + 12 more
Today’s science increasingly requires effective ways to find and access existing datasets that are distributed across a range of repositories. For researchers in the life sciences, discoverability of datasets may soon become as essential as identifying the latest publications via PubMed. Through an international…
Sarah Ciston, Mike Ananny, Kate Crawford
| 1 | | --- | | INTRODUCTION TO MACHINE LEARNING | | DATASETS | | Maybe you're an engineer creating a new machine vision system to track birds. You | | might be a journalist using social media data to research Costa Rican households. You could be a researcher who stumbled upon your university's archive of | |…
Authors not listed
The field of computational chemistry is increasingly leveraging machine learning (ML) potentials to predict molecular properties with high accuracy and efficiency, providing a viable alternative to traditional quantum mechanical (QM) methods, which are often computationally intensive. Central to the success of ML…
Paul Bilokon, Oleksandr Bilokon, Saeed Amen
Recent advances in data science, machine learning, and artificial intelligence, such as the emergence of large language models, are leading to an increasing demand for data that can be processed by such models. While data sources are application-specific, and it is impossible to produce an exhaustive list of such data…
Ryan Quey, Matthew A. Schiefer, Anmol Kiran, Bhavesh Patel
This manuscript provides the methods and outcomes of KnowMore, the Grand Prize winning automated knowledge discovery tool developed by our team during the 2021 NIH SPARC FAIR Data Codeathon. The National Institutes of Health Stimulating Peripheral Activity to Relieve Conditions (NIH SPARC) program generates rich…
Tjelvar S.G. Olsson, Matthew Hartley, Todd Vision
The explosion in volumes and types of data has led to substantial challenges in data management. These challenges are often faced by front-line researchers who are already dealing with rapidly changing technologies and have limited time to devote to data management. There are good high-level guidelines for managing and…
Grace S. Brown, James Wengler, Aaron Joyce S. Fabelico, Abigail Muir + 7 more
Millions of high-throughput, molecular datasets have been shared in public repositories. have been shared in public repositories. Researchers can reuse such data to validate their own findings and explore novel questions. A frequent goal is to find multiple datasets that address similar research topics and to either…
Omar Benjelloun, Shiyu Chen, Natasha Noy
Scientists, governments, and companies increasingly publish datasets on the Web. Google's Dataset Search extracts dataset metadata expressed using schema.org and similar vocabularies—from Web pages in order to make datasets discoverable. Since we started the work on Dataset Search in 2016, the number of datasets…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…
Fritz Lekschas, Nils Gehlenborg
The number of data sets in biomedical repositories has grown rapidly over the past decade, providing scientists in fields like genomics and other areas of high-throughput biology with tremendous opportunities to re-use data. Scientists are able to test hypotheses computationally instead of generating their own data, to…
John Kratz, Carly Strasser
The movement to bring datasets into the scholarly record as first class research products (validated, preserved, cited, and credited) has been inching forward for some time, but now the pace is quickening. As data publication venues proliferate, significant debate continues over formats, processes, and terminology.…
Evan Walter Clark Spotte-Smith, Orion Archer Cohen, Samuel Blau, Jason Munro + 8 more
Advanced chemical research is increasingly reliant on large computed datasets of molecules and reactions to discover new functional molecules, understand chemicaltrends,train machine learning models, and more. To be of greatest use to the scientific community, such datasets should follow FAIR principles (i.e. be…
Pascal Petit, Nicolas Vuillerme
Data have become central to scientific discovery. While primary data collection remains vital, there is growing recognition of the benefits of reusing existing datasets. However, identifying suitable datasets for specific research questions is increasingly difficult due to the fragmentation and heterogeneity of the big…
Manuel Blázquez Ochando, Juan José Prieto Gutiérrez
This paper presents an analysis of the publication of datasets collected via Google Dataset Search, specialized in families of RNA viruses, whose terminology was obtained from the National Cancer Institute (NCI) thesaurus developed by the US Department of Health and Human Services. The objective is to determine the…
Alisa Bokulich, Wendy Parker
We critically engage two traditional views of scientific data and outline a novel philosophical view that we call the pragmatic-representational (PR) view of data. On the PR view, data are representations that are the product of a process of inquiry, and they should be evaluated in terms of their adequacy or fitness…
Ramanathan V. Guha, Prashanth Radhakrishnan, Bo Xu, Wei Sun + 9 more
'Carolyn Au' 'Ajai Tirumali' 'Muhammad Jehangir Amjad' 'Samantha N. Piekos' 'Natalie Diaz' 'Jennifer Chen' 'Julia Wu' 'Prem Ramaswami' 'James Manyika'] Publicly available data from open sources (e.g., United States Census Bureau (Census) [1], World Health Organization (WHO) [2], Intergovernmental Panel on Climate…
Alfredo Nazábal, Christopher K. I. Williams, Giovanni Colavizza, Camila Rangel Smith + 1 more
'Camila Rangel Smith' 'A. A. Williams'] Consider the situation where a data analyst wishes to carry out an analysis on a given dataset. It is widely recognized that most of the analyst's time will be taken up with data engineering tasks such as acquiring, understanding, cleaning and preparing the data. In this paper we…
David Buterez, Jon Paul Janet, Steven J. Kiddle, Pietro Liò
High-throughput screening (HTS), as one of the key techniques in drug discovery, is frequently used to identify promising drug candidates in a largely automated and cost-effective way. One of the necessary conditions for successful HTS campaigns is a large and diverse compound library, enabling hundreds of thousands of…
Authors not listed
The discoverability and reusability of data is critical for machine learning to drive new discovery in the chemical sciences, and the ‘FAIR Guiding Principles for scientific data management and stewardship’ provide a measurable set of guidelines that can be used to ensure the accessibility of reusable data. We…
George Alter, Alejandra Gonzalez-Beltran, Lucila Ohno-Machado, Philippe Rocca-Serra
This article presents elements in the Data Tags Suite (DATS) metadata schema describing data access, data use conditions, and consent information. DATS is a product of the bioCADDIE Project, which created a data discovery index for searching across all types of biomedical data. The “access and use” metadata items in…
Mike A. Thelwall, Marcus Munafò, Amalia Mas Bleda, Emma Stuart + 6 more
Primary data collected during a research study is increasingly shared and may be re-used for new studies. To assess the extent of data sharing in favourable circumstances and whether such checks can be automated, this article investigates the summary statistics of primary human genome-wide association studies (GWAS).…
David J. Hand
Ready data availability, cheap storage capacity, and powerful tools for extracting information from data have the potential to significantly enhance the human condition. However, as with all advanced technologies, this comes with the potential for misuse. Ethical oversight and constraints are needed to ensure that an…
David Buterez, Jon Paul Janet, Steven Kiddle, Pietro Liò
High throughput screening (HTS) is one of the leading techniques for hit identification in drug discovery and comprises of multiple phases, one primary and one or more confirmatory screens which result in multi-fidelity data. Noisy primary screening data are available on a large number of compounds and higher quality…
Authors not listed
Raman spectroscopy is an increasingly powerful and fast-growing analytical technique across diverse disciplines, from materials science and chemistry to biology and medicine, thanks to advances in Raman instrumentation and greatly supported by the flourishing of chemometrics and artificial intelligence (AI). However…