26 papers · ranked by Valyu relevance
Lucina Hackman, Pauline Mack, Hervé Ménard
Data underpinning science have become one of the most precious assets in research, and while the principles of FAIR (Findable, Accessible, Interoperable and Reusable) have been put forward as a guide to how to approach data handling, data sharing and long-term storage still remain a challenge for many research areas…
David J. Hand
Ready data availability, cheap storage capacity, and powerful tools for extracting information from data have the potential to significantly enhance the human condition. However, as with all advanced technologies, this comes with the potential for misuse. Ethical oversight and constraints are needed to ensure that an…
Claudio Gutiérrez
The foundations of experience (since we absolutely must get down to this) have been non-existent or very weak; nor has a collection or store of particulars yet been sought or made, able or in any way adequate, either in number, kind or certainty, to inform the intellect. [...] Natural history contains nothing that has…
Ian L. Boyd
Comment When I think of data I think of binary or hexadecimal numbers. This betrays something of my background, but it was a surprise to me when in Defra, the UK Department of State with responsibility for food and the environment, we started to talk about data and I found that other people saw data very differently.…
Dimitri Yatsenko, Jacob Reimer, Alexander S. Ecker, Edgar Y. Walker + 6 more
The rise of big data in modern research poses serious challenges for data management: Large and intricate datasets from diverse instrumentation must be precisely aligned, annotated, and processed in a variety of ways to extract new insights. While high levels of data integrity are expected, research teams have diverse…
Roman Lukyanenko
Data Management Authors: ['Roman Lukyanenko'] In an era dominated by information technology, the critical discipline of data management remains undervalued compared to the innovations it enables, such as artificial intelligence and social media. The ambiguity surrounding what constitutes data management and its…
M. Zanin, D. Papo, P. A. Sousa, E. Menasalvas + 3 more
The increasing power of computer technology does not dispense with the need to extract meaningful in-formation out of data sets of ever growing size, and indeed typically exacerbates the complexity of this task. To tackle this general problem, two methods have emerged, at chronologically different times, that are now…
Petar Radanliev, David De Roure
With the increased digitalisation of our society, new and emerging forms of data present new values and opportunities for improved data driven multimedia services, or even new solutions for managing future global pandemics (i.e., Disease X). This article conducts a literature review and bibliometric analysis of…
M. TAMER ÖZSU
Work-in-progressThere has been an increasing recognition of the value of data and of data-based decision making. As a consequence, the development of data science as a field of study has intensified in recent years. However, there is no systematic and comprehensive treatment and understanding of data science. This…
Rafael C. Alvarado
Consensus on the definition of data science remains low despite the widespread establishment of academic programs in the field and continued demand for data scientists in industry. Definitions range from rebranded statistics to data-driven science to the science of data to simply the application of machine learning to…
Gianluca Brunori, Manlio Bacco, Carolina Puerta-Piñero, Maria Teresa Borzacchiello + 1 more
'Maria Teresa Borzacchiello' 'Eckhard Stormer'] Title: Highlights 1. • The European Union (EU) is investing in developing Common European Data Spaces in several domains, including agriculture, pushing for a vibrant data market and data exploitation in the years to come. 2. • The expected benefits for the farmers…
S Bauermeister, C Orton, S Thompson, R A Barker + 62 more
The Dementias Platform UK (DPUK) Data Portal is a data repository facilitating access to data for 3 370 929 individuals in 42 cohorts. The Data Portal is an end-to-end data management solution providing a secure, fully auditable, remote access environment for the analysis of cohort data. All projects utilising the data…
Matthias Scheffler, Stefan Bauer, Peter Benner, Tristan Bereau + 57 more
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Bohdan B. Khomtchouk, Kasra A. Vand, Thor Wahlestedt, Kelly Khomtchouk + 2 more
We propose a search engine and file retrieval system for all bioinformatics databases worldwide. PubData searches biomedical data in a user-friendly fashion similar to how PubMed searches biomedical literature. PubData is built on novel network programming, natural language processing, and artificial intelligence…
Dominik Balazka, Dario Rodighiero
Starting from an analysis of frequently employed definitions of big data, it will be argued that, to overcome the intrinsic weaknesses of big data, it is more appropriate to define the object in relational terms. The excessive emphasis on volume and technological aspects of big data, derived from their current…
W. Anderson, R. Apweiler, A. Bateman, G.A. Bauer + 27 more
On November 18-19, 2016, the Human Frontier Science Program Organization (HFSPO) hosted a meeting of senior managers of key data resources and leaders of several major funding organizations to discuss the challenges associated with sustaining biological and biomedical (i.e., life sciences) data resources and associated…
Susanna-Assunta Sansone, Alejandra Gonzalez-Beltran, Philippe Rocca-Serra, George Alter + 12 more
Today’s science increasingly requires effective ways to find and access existing datasets that are distributed across a range of repositories. For researchers in the life sciences, discoverability of datasets may soon become as essential as identifying the latest publications via PubMed. Through an international…
Authors not listed
The discoverability and reusability of data is critical for machine learning to drive new discovery in the chemical sciences, and the ‘FAIR Guiding Principles for scientific data management and stewardship’ provide a measurable set of guidelines that can be used to ensure the accessibility of reusable data. We…
Longbing Cao
—Data science is creating very exciting trends as well as significant controversy. A critical matter for the healthy development of data science in its early stages is to deeply understand the nature of data and data science, and to discuss the various pitfalls. These important issues motivate the discussions in this…
Matthias Scheffler
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Rebecca Grant, Iain Hrynaszkiewicz
This paper describes the adoption of a standard policy for the inclusion of data availability statements in all research articles published at the Nature family of journals, and the subsequent research which assessed the impacts that these policies had on authors, editors, and the availability of datasets. The key…
Paul R. Burton, Madeleine J. Murtagh, Andy Boyd, James B. Williams + 12 more
'Edward S. Dove' 'Susan E. Wallace' 'Anne-Marie Tassé' 'Julian Little' 'Rex L. Chisholm' 'Amadou Gaye' 'Kristian Hveem' 'Anthony J. Brookes' 'Pat Goodwin' 'Jon Fistein' 'Martin Bobrow' 'Bartha M. Knoppers'] Motivation: The data that put the ‘evidence’ into ‘evidence-based medicine’ are central to developments in public…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Rebecca Brunk, Kriti Shukla, Bryant Hutson, Yue Wang + 7 more
Genomic sequencing and other big biological data is unquestionably of paramount value, however the success in recruiting highly skilled individuals with diverse backgrounds has been limited. A main reason for this deficiency could be due to the lack of educational resources and early exposure to the field. With the…
Aaron Liu, Myeongyeon Lee, Rahul Venkatesh, Jessica Bonsu + 4 more
Polymer-based semiconductors and organic electronics encapsulate a significant research thrust for informatics-driven materials development. However, device measurements are described by a complex array of design and parameter choices, many of which are sparsely reported. For example, the mobility of a polymer-based…