24 papers · ranked by Valyu relevance
Dani Arribas-Bel, Mark Green, Francisco Rowe, Alex Singleton
This paper develops the notion of “open data product”. We define an open data product as the open result of the processes through which a variety of data (open and not) are turned into accessible information through a service, infrastructure, analytics or a combination of all of them, where each step of development is…
Lucina Hackman, Pauline Mack, Hervé Ménard
Data underpinning science have become one of the most precious assets in research, and while the principles of FAIR (Findable, Accessible, Interoperable and Reusable) have been put forward as a guide to how to approach data handling, data sharing and long-term storage still remain a challenge for many research areas…
Li, Zhi, Zhang, Lei + 10 more
The data circulation is a complex scenario involving a large number of participants and different types of requirements, which not only has to comply with the laws and regulations, but also faces multiple challenges in technical and business areas. In order to systematically and comprehensively address these issues, it…
Blend Berisha, Endrit Mëziu, Isak Shabani
Big Data and Cloud Computing as two mainstream technologies, are at the center of concern in the IT field. Every day a huge amount of data is produced from different sources. This data is so big in size that traditional processing tools are unable to deal with them. Besides being big, this data moves fast and has a lot…
Teresa Gomez-Diaz, Tomas Recio
Background: Research Software is a concept that has been only recently clarified. In this paper we address the need for a similar enlightenment concerning the Research Data concept. Methods: Our contribution begins by reviewing the Research Software definition, which includes the analysis of software as a legal…
Isaac Virshup, Sergei Rybakov, Fabian J. Theis, Philipp Angerer + 1 more
anndata is a Python package for handling annotated data matrices in memory and on disk (github.com/theislab/anndata), positioned between pandas and xarray. anndata offers a broad range of computationally efficient features including, among others, sparse data support, lazy operations, and a PyTorch interface.…
Roman Lukyanenko
Data Management Authors: ['Roman Lukyanenko'] In an era dominated by information technology, the critical discipline of data management remains undervalued compared to the innovations it enables, such as artificial intelligence and social media. The ambiguity surrounding what constitutes data management and its…
L. Clissa, M. Lassnig, L. Rinaldi
The modern increase in data production is driven by multiple factors, and several stakeholders from various sectors contribute to it. Although drawing a comparison of the sizes at stake for different big data players is hard due to the lack of official data, this report tries to reconstruct the yearly orders of…
Adina S. Wagner, Laura K. Waite, Małgorzata Wierzba, Felix Hoffstaedter + 4 more
Large-scale datasets present unique opportunities to perform scientific investigations with unprecedented breadth. However, they also pose considerable challenges for the findability, accessibility, interoperability, and reusability (FAIR) of research outcomes due to infrastructure limitations, data usage constraints…
Petar Radanliev, David De Roure
With the increased digitalisation of our society, new and emerging forms of data present new values and opportunities for improved data driven multimedia services, or even new solutions for managing future global pandemics (i.e., Disease X). This article conducts a literature review and bibliometric analysis of…
Gianluca Brunori, Manlio Bacco, Carolina Puerta-Piñero, Maria Teresa Borzacchiello + 1 more
'Maria Teresa Borzacchiello' 'Eckhard Stormer'] Title: Highlights 1. • The European Union (EU) is investing in developing Common European Data Spaces in several domains, including agriculture, pushing for a vibrant data market and data exploitation in the years to come. 2. • The expected benefits for the farmers…
M. TAMER ÖZSU
Work-in-progressThere has been an increasing recognition of the value of data and of data-based decision making. As a consequence, the development of data science as a field of study has intensified in recent years. However, there is no systematic and comprehensive treatment and understanding of data science. This…
Rafael C. Alvarado
Consensus on the definition of data science remains low despite the widespread establishment of academic programs in the field and continued demand for data scientists in industry. Definitions range from rebranded statistics to data-driven science to the science of data to simply the application of machine learning to…
Matthias Scheffler, Stefan Bauer, Peter Benner, Tristan Bereau + 57 more
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Connor Bernard, Gabriel Silva Santos, Jacques Deere, Roberto Rodriguez-Caro + 5 more
The ecological sciences have joined the big data revolution. However, despite exponential growth in data availability, broader interoperability amongst datasets is still needed to unlock the potential of open access. The interface of demography and functional traits is well-positioned to benefit from said…
Authors not listed
The discoverability and reusability of data is critical for machine learning to drive new discovery in the chemical sciences, and the ‘FAIR Guiding Principles for scientific data management and stewardship’ provide a measurable set of guidelines that can be used to ensure the accessibility of reusable data. We…
Tobias K. Mildenberger, Federico Maioli, Casper W. Berg
Scientific bottom-trawl surveys provide essential fisheries-independent data for fisheries and ecosystem research. In the Northeast Atlantic, the ICES Database of Trawl Surveys (DATRAS) compiles haul-level information, species- and length-specific catch data, and individual biological observations across multiple…
Matthias Scheffler
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
J.L. Hyde, A.C. Swanson, S.A. Bohlman, S. Athayde + 2 more
New infrastructure projects are planned or under construction in several countries, including in the bioculturally diverse Amazon, Mekong, and Congo regions. While infrastructure development can improve human health and living standards, it may also lead to environmental degradation and social change. Accessible, high…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Manuel Schottdorf, Guoqiang Yu, Edgar Y. Walker
The rise of large scientific collaborations in neuroscience requires systematic, scalable, and reliable data management. How this is best done in practice remains an open question. To address this, we conducted a data science survey among currently active U19 grants, funded through the NIH’s BRAIN Initiative. The…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Rebecca Brunk, Kriti Shukla, Bryant Hutson, Yue Wang + 7 more
Genomic sequencing and other big biological data is unquestionably of paramount value, however the success in recruiting highly skilled individuals with diverse backgrounds has been limited. A main reason for this deficiency could be due to the lack of educational resources and early exposure to the field. With the…
Aaron Liu, Myeongyeon Lee, Rahul Venkatesh, Jessica Bonsu + 4 more
Polymer-based semiconductors and organic electronics encapsulate a significant research thrust for informatics-driven materials development. However, device measurements are described by a complex array of design and parameter choices, many of which are sparsely reported. For example, the mobility of a polymer-based…