23 papers · ranked by Valyu relevance
Amaryllis Mavragani, Ahmed Hassan, Tushar Khinvasara, AasimAyaz Wani + 4 more
'Anthony Lighterness' 'Michael Adcock' 'Lauren Abigail Scanlon' 'Gareth Price'] Background The promise of real-world evidence and the learning health care system primarily depends on access to high-quality data. Despite widespread awareness of the prevalence and potential impacts of poor data quality (DQ), best…
Rainer Tinscher Thorsten Wuest
In materials sciences, a large amount of research data is generated through a broad spectrum of different experiments. As of today, experimental research data including meta-data in materials science is often stored decentralized by the researcher(s) conducting the experiments without generally accepted standards on…
Allan Koch Veiga, Antonio Mauro Saraiva, Arthur David Chapman, Paul John Morris + 4 more
'Paul John Morris' 'Christian Gendreau' 'Dmitry Schigel' 'Tim James Robertson' 'Ulrich Melcher'] The increasing availability of digitized biodiversity data worldwide, provided by an increasing number of institutions and researchers, and the growing use of those data for a variety of purposes have raised concerns…
Naila A. Shaheen, Bipin Manezhi, Abin Thomas, Mohammed AlKelya
Background A dataset is indispensable to answer the research questions of clinical research studies. Inaccurate data lead to ambiguous results, and the removal of errors results in increased cost. The aim of this Quality Improvement Project (QIP) was to improve the Data Quality (DQ) by enhancing conformance and…
Authors not listed
Raman spectroscopy is an increasingly powerful and fast-growing analytical technique across diverse disciplines, from materials science and chemistry to biology and medicine, thanks to advances in Raman instrumentation and greatly supported by the flourishing of chemometrics and artificial intelligence (AI). However…
Sijie Dong, Soror Sahri, Themis Palpanas
Data Science Systems Authors: ['Sijie Dong' 'Soror Sahri' 'Themis Palpanas'] Artificial intelligence (AI) has transformed various fields, significantly impacting our daily lives. A major factor in AI's success is high-quality data. In this paper, we present a comprehensive review of the evolution of data quality (DQ)…
Christian Lovis, Delphine Courvoisier, Zhan Wang, Jens Declerck + 3 more
Background Health care has not reached the full potential of the secondary use of health data because of-among other issues-concerns about the quality of the data being used. The shift toward digital health has led to an increase in the volume of health data. However, this increase in quantity has not been matched by a…
Sonja Harkener, Oliver J Bott, Christian Draeger, Tobias Hartz + 6 more
Background To be beneficial for empirical health research, a dataset must be fit for use. The quality of a dataset can only be influenced during data collection, yet it is evaluated multiple times during analysis or secondary use by applying quality indicators. Objective This study aimed to establish an up-to-date set…
Heidi Carolina Tamm, Anastasija Nikiforova
Rule Definition in Data Warehouses Authors: ['Heidi Carolina Tamm' 'Anastasija Nikiforova'] Abstract: In the contemporary data-driven landscape, ensuring data quality (DQ) is crucial for deriving actionable insights from vast data repositories. The objective of this study is to explore the potential for automating data…
Anastasija Nikiforova
Nowadays open data is entering the mainstream - it is free available for every stakeholder and is often used in business decision-making. It is important to be sure data is trustable and error-free as its quality problems can lead to huge losses. The research discusses how (open) data quality could be assessed. It also…
Tiffany Leung, Marianna Kapsetaki, Ibrahim Adeleke, Hareesh Veldandi + 8 more
'W. David Dotson' 'Ting Fang Alvin Ang' 'Filipe Andrade Bernardi' 'Domingos Alves' 'Nathalia Crepaldi' 'Diego Bettiol Yamada' 'Vinícius Costa Lima' 'Rui Rijo'] Background Decision-making and strategies to improve service delivery must be supported by reliable health data to generate consistent evidence on health…
Fernando Gualo, Moisés Rodríguez, Javier Verdugo, Ismael Caballero + 1 more
'Mario Piattini'] > Abstract. The most successful organizations in the world are data-driven businesses. Data is at the core of the business of many organizations as one of the most important assets, since the decisions they make cannot be better than the data on which they are based. Due to this reason, organizations…
Margaret Kosmala, Andrea Wiggins, Alexandra Swanson, Brooke Simmons
Ecological and environmental citizen science projects have enormous potential to advance science, influence policy, and guide resource management by producing datasets that are otherwise infeasible to generate. This potential can only be realized, though, if the datasets are of high quality. While scientists are often…
Qingyu Chen, Ramona Britto, Ivan Erill, Constance J. Jeffery + 7 more
The volume of biological database records is growing rapidly, populated by complex records drawn from heterogeneous sources. A specific challenge is duplication, that is, the presence of redundancy (records with high similarity) or inconsistency (dissimilar records that correspond to the same entity). The…
Eric W. Deutsch, Roger Kramer, Joseph Ames, Andrew Bauman + 21 more
Translational biomedical research is generating exponentially more data: thousands of whole-genome sequences (WGS) are now available; brain data are doubling every two years. Analyses of Big Data, including imaging, genomic, phenotypic, and clinical data, present qualitatively new challenges as well as opportunities.…
Authors not listed
A Python script for the systematic, high-throughput analysis of accurate mass data was developed and tested on over 3,000 Supporting Information (SI) PDFs from Organic Letters. For each SI file, quadruplets of molecular formula, measured ion, e.g. [M+Na]+, reported calculated and found masses were extracted and…
Yu-Chieh Huang, Pierre Tremouilhac, Stefan Kuhn, Pei-Chi Huang + 6 more
A method for data review in chemical sciences with a focus on data for the characterization of synthetic molecules is described. As current procedures for data curation in chemistry rely almost exclusively on manual checking or peer reviewing, a (semi-)automatic procedure for the evaluation of data assigned to…
IC Kos-Braun, B Gerlach, C Pitzer
Recently, it has become evident that academic research faces issues with the reproducibility of research data. It is critical to understand the underlying causes in order to remedy this situation. Core Facilities (CFs) have a central position in the research infrastructure and therefore they are ideally suited to…
Authors not listed
The precision of thermodynamic modeling for ionic liquid (IL)–solute systems is fundamentally reliant on the quality of experimental data. However, prevalent databases such as ILThermo frequently exhibit conflicting measurements for the same systems under identical temperature and pressure conditions. These disparities…
Glenda M. Yenni, Erica M. Christensen, Ellen K. Bledsoe, Sarah R. Supp + 3 more
Data management and publication are core components of the research process. An emerging challenge that has received limited attention in biology is managing, working with, and providing access to data under continual active collection. “Evolving data” present unique challenges in quality assurance and control, data…
Rebecca Grant, Graham Smith, Iain Hrynaszkiewicz
Since 2017, the publisher Springer Nature has provided an optional Research Data Support service to help researchers deposit and curate data that support their peer-reviewed publications. This service builds on a Research Data Helpdesk, which since 2016 has provided support to authors and editors who need advice on the…
Ammar Ammar, Chris Evelo, Egon Willighagen
New nanomaterials improve our society. Understanding their effects on biological systems is of importance to improve our understanding of their properties and safety. However, reusability of previously produced data to help developing computational risk assessment tools is still limited, due to the inconsistency in…
Authors not listed
Metabolomics studies require complex data processing pipelines to ensure data quality and extract meaningful biological insights. GetFeatistics is an R-package developed to streamline the elaboration and statistical analysis of metabolomics data. For targeted analyses, the package enables calibration curve-based…