24 papers · ranked by Valyu relevance
Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis + 3 more
'Luis Ibáñez' 'Emilia Kacprzak' 'Paul Groth'] Abstract Generating value from data requires the ability to find, access and make sense of datasets. There are many efforts underway to encourage data sharing and reuse, from scientific publishers asking authors to submit data alongside manuscripts to data marketplaces…
Lucina Hackman, Pauline Mack, Hervé Ménard
Data underpinning science have become one of the most precious assets in research, and while the principles of FAIR (Findable, Accessible, Interoperable and Reusable) have been put forward as a guide to how to approach data handling, data sharing and long-term storage still remain a challenge for many research areas…
Michel Dumontier, Alasdair J.G. Gray, M. Scott Marshall, Vladimir Alexiev + 27 more
'Vladimir Alexiev' 'Peter Ansell' 'Gary Bader' 'Joachim Baran' 'Jerven T. Bolleman' 'Alison Callahan' 'José Cruz-Toledo' 'Pascale Gaudet' 'Erich A. Gombocz' 'Alejandra N. Gonzalez-Beltran' 'Paul Groth' 'Melissa Haendel' 'Maori Ito' 'Simon Jupp' 'Nick Juty' 'Toshiaki Katayama' 'Norio Kobayashi' 'Kalpana Krishnaswami'…
Leigh Dodds
The FAIR principles need to be applied in context. To do that, we need to understand both the needs of data users and the characteristics of the data to be shared. This Opinion introduces ten different dataset archetypes that can be used to inform plans for how data are to be accessed, used, and shared.
John Kratz, Carly Strasser
The movement to bring datasets into the scholarly record as first class research products (validated, preserved, cited, and credited) has been inching forward for some time, but now the pace is quickening. As data publication venues proliferate, significant debate continues over formats, processes, and terminology.…
Omar Benjelloun, Shiyu Chen, Natasha Noy
Scientists, governments, and companies increasingly publish datasets on the Web. Google's Dataset Search extracts dataset metadata expressed using schema.org and similar vocabularies—from Web pages in order to make datasets discoverable. Since we started the work on Dataset Search in 2016, the number of datasets…
Authors not listed
The field of computational chemistry is increasingly leveraging machine learning (ML) potentials to predict molecular properties with high accuracy and efficiency, providing a viable alternative to traditional quantum mechanical (QM) methods, which are often computationally intensive. Central to the success of ML…
Fritz Lekschas, Nils Gehlenborg
The number of data sets in biomedical repositories has grown rapidly over the past decade, providing scientists in fields like genomics and other areas of high-throughput biology with tremendous opportunities to re-use data. Scientists are able to test hypotheses computationally instead of generating their own data, to…
Paul Bilokon, Oleksandr Bilokon, Saeed Amen
Recent advances in data science, machine learning, and artificial intelligence, such as the emergence of large language models, are leading to an increasing demand for data that can be processed by such models. While data sources are application-specific, and it is impossible to produce an exhaustive list of such data…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…
Behnam Ghavimi, Philipp Mayr, Christoph Lange, Sahar Vahdati + 1 more
'Sören Auer'] > Abstract. Today, full-texts of scientific articles are often stored in different locations than the used datasets. Dataset registries aim at a closer integration by making datasets citable but authors typically refer to datasets using inconsistent abbreviations and heterogeneous metadata (e.g. title…
Manuel Blázquez Ochando, Juan José Prieto Gutiérrez
This paper presents an analysis of the publication of datasets collected via Google Dataset Search, specialized in families of RNA viruses, whose terminology was obtained from the National Cancer Institute (NCI) thesaurus developed by the US Department of Health and Human Services. The objective is to determine the…
Alisa Bokulich, Wendy Parker
We critically engage two traditional views of scientific data and outline a novel philosophical view that we call the pragmatic-representational (PR) view of data. On the PR view, data are representations that are the product of a process of inquiry, and they should be evaluated in terms of their adequacy or fitness…
Alfredo Nazábal, Christopher K. I. Williams, Giovanni Colavizza, Camila Rangel Smith + 1 more
'Camila Rangel Smith' 'A. A. Williams'] Consider the situation where a data analyst wishes to carry out an analysis on a given dataset. It is widely recognized that most of the analyst's time will be taken up with data engineering tasks such as acquiring, understanding, cleaning and preparing the data. In this paper we…
Helle W. van den Maagdenberg, Martin Šícho, David Alencar Araripe, Sohvi Luukkonen + 9 more
Building reliable and robust quantitative structure-property relationship (QSPR) models is a challenging task. First, the experimental data needs to be obtained, analyzed and curated. Second, the number of available methods is continuously growing and evaluating different algorithms and methodologies can be arduous.…
Simon Boothroyd, Lee-Ping Wang, David Mobley, John Chodera + 1 more
Developing accurate classical force field representations of molecules is key to realizing the full potential of molecular simulations, both as a powerful route to gaining fundamental insight into a broad spectrum of chemical and biological phenomena, and for predicting physicochemical and mechanical properties of…
Benjamin D. Morris, Ethan P. White, Nathan G. Swenson
Ecological research relies increasingly on the use of previously collected data. Use of existing datasets allows questions to be addressed more quickly, more generally, and at larger scales than would otherwise be possible. As a result of large-scale data collection efforts, and an increasing emphasis on data…
Rebecca Grant, Iain Hrynaszkiewicz
This paper describes the adoption of a standard policy for the inclusion of data availability statements in all research articles published at the Nature family of journals, and the subsequent research which assessed the impacts that these policies had on authors, editors, and the availability of datasets. The key…
Dimitri Yatsenko, Jacob Reimer, Alexander S. Ecker, Edgar Y. Walker + 6 more
The rise of big data in modern research poses serious challenges for data management: Large and intricate datasets from diverse instrumentation must be precisely aligned, annotated, and processed in a variety of ways to extract new insights. While high levels of data integrity are expected, research teams have diverse…
Connor Bernard, Gabriel Silva Santos, Jacques Deere, Roberto Rodriguez-Caro + 5 more
The ecological sciences have joined the big data revolution. However, despite exponential growth in data availability, broader interoperability amongst datasets is still needed to unlock the potential of open access. The interface of demography and functional traits is well-positioned to benefit from said…
Elizabeth Wenk, Payal Bal, David Coleman, Rachael Gallagher + 2 more
Trait databases have proliferated over the past decades, facilitating research on the ecology, evolution, and conservation of taxa across the Tree of Life. Typically, teams of independent researchers build these databases, and each must develop their own workflow and output structure. This divests research hours from…
Aaron Liu, Myeongyeon Lee, Rahul Venkatesh, Jessica Bonsu + 4 more
Polymer-based semiconductors and organic electronics encapsulate a significant research thrust for informatics-driven materials development. However, device measurements are described by a complex array of design and parameter choices, many of which are sparsely reported. For example, the mobility of a polymer-based…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Yovaninna Alarcón‐Soto, Jenifer Espasandín-Domínguez, Ipek Guler, Mercedes Conde‐Amboage + 4 more
'Mercedes Conde‐Amboage' 'Francisco Gudé' 'Klaus Langohr' 'Carmén Cadarso-Suárez' 'Guadalupe Gómez Melis'] Abstract: We highlight the role of Data Science in Biomedicine. Our manuscript goes from the general to the particular, presenting a global definition of Data Science and showing the trend for this discipline…