23 papers · ranked by Valyu relevance
Cindy Cheng, Luca Messerschmidt, Isaac Bravo, Marco Waldbauer + 6 more
Data harmonization is an important method for combining or transforming data. To date however, articles about data harmonization are field-specific and highly technical, making it difficult for researchers to derive general principles for how to engage in and contextualize data harmonization efforts. This commentary…
Jimmy K. Yu, Marcos Martínez-Romero, Matthew Horridge, Mete U. Akdogan + 1 more
'Mete U. Akdogan' 'Mark A. Musen'] In the age of big data, it is important for primary research data to follow the FAIR principles of findability, accessibility, interoperability, and reusability. Data harmonization enhances interoperability and reusability by aligning heterogeneous data under standardized…
Ivan S. Gill, Emma J. Griffiths, Damion Dooley, Rhiannon Cameron + 31 more
'Sarah Savić Kallesøe' 'Nithu Sara John' 'Anoosha Sehar' 'Gurinder Gosal' 'David Alexander' 'Madison Chapel' 'Matthew A. Croxen' 'Benjamin Delisle' 'Rachelle Di Tullio' 'Daniel Gaston' 'Ana Duggan' 'Jennifer L. Guthrie' 'Mark Horsman' 'Esha Joshi' 'Levon Kearny' 'Natalie Knox' 'Lynette Lau' 'Jason J. LeBlanc' 'Vincent…
Aécio Santos, Eduardo H. M. Pena, Roque López, Juliana Freire
Data harmonization is an essential task that entails integrating datasets from diverse sources. Despite years of research in this area, it remains a time-consuming and challenging task due to schema mismatches, varying terminologies, and differences in data collection methodologies. This paper presents the case for…
Bey-Marrié Schmidt, Christopher J. Colvin, Ameer Hohlfeld, Natalie Leon
'Natalie Leon'] Background Data harmonisation is an important intervention to strengthen health systems functioning. It has the potential to enhance the production, accessibility and utilisation of routine health information for clinical and service management decision-making. It is important to understand the range of…
Kaelyn Long, Kai Gravel-Pucillo, Levi Waldron, Sean Davis + 1 more
Public omics repositories contain vast amounts of valuable data, but their metadata suffers from extreme heterogeneity, unstandardized terminologies, and quality issues that severely limit data reusability and cross-study integration. While prospective metadata standards exist, the majority of published omics data…
Joeri Kalter, Maike G. Sweegers, Irma M. Verdonck-de Leeuw, Johannes Brug + 1 more
Objective Harmonizing individual patient data (IPD) for meta-analysis has clinical and statistical advantages. Harmonizing IPD from multiple studies may benefit from a flexible data harmonization platform (DHP) that allows harmonization of IPD already during data collection. This paper describes the development and use…
Seyed Amir Tabatabaei Hosseini, Reza Kazemzadeh, Bethany Joy Foster, Emre Arpali + 1 more
'Emre Arpali' 'Caner Süsal'] In organ transplantation, accurate analysis of clinical outcomes requires large, high-quality data sets. Not only are outcomes influenced by a multitude of factors such as donor, recipient, and transplant characteristics and posttransplant events but they may also change over time. Although…
Bey-Marrié Schmidt, Christopher J. Colvin, Ameer Hohlfeld, Natalie Leon
'Natalie Leon'] Background Data harmonisation (DH) has emerged amongst health managers, information technology specialists and researchers as an important intervention for routine health information systems (RHISs). It is important to understand what DH is, how it is defined and conceptualised, and how it can lead to…
Steven Wilkins‐Reeves, Yen‐Chi Chen, Kwun Chuen Gary Chan
Data harmonization is the process by which an equivalence is developed between two variables measuring a common trait. Our problem is motivated by dementia research in which multiple tests are used in practice to measure the same underlying cognitive ability such as language or memory. We connect this statistical…
Adrienne M. Stilp, Leslie S. Emery, Jai G. Broome, Erin J. Buth + 68 more
Genotype-phenotype association studies often combine phenotype data from multiple studies to increase power. Harmonization of the data usually requires substantial effort due to heterogeneity in phenotype definitions, study design, data collection procedures, and data set organization. Here we describe a centralized…
Matthew Pearce, Tom R.P. Bishop, Stephen Sharp, Kate Westgate + 3 more
Harmonisation of data for pooled analysis relies on the principle of inferential equivalence between variables from different sources. Ideally, this is achieved using models of the direct relationship with gold standard criterion measures, but the necessary validation data are often unavailable. This study examines an…
Heming Zhang, Shunning Liang, Tim Xu, Wenyu Li + 15 more
Artificial intelligence (AI) is revolutionizing scientific discovery because of its super capability, following the neural scaling laws, to integrate and analyze large-scale datasets to mine knowledge. Foundation models, large language models (LLMs) and large vision models (LVMs), are among the most important…
Authors not listed
Raman spectroscopy is an increasingly powerful and fast-growing analytical technique across diverse disciplines, from materials science and chemistry to biology and medicine, thanks to advances in Raman instrumentation and greatly supported by the flourishing of chemometrics and artificial intelligence (AI). However…
Niek F. de Jonge, Helge Hecht, Justin J. J. van der Hooft, Florian Huber
Mass spectral libraries have proven to be essential for mass spectrum annotation, both for library matching and training new machine learning algorithms. A key step in training machine learning models is having high-quality training data. Public libraries of mass spectrometry data that are open to user submission often…
Mark D. Danese, Marc Halperin, Jennifer Duryea, Ryan Duryea
Most healthcare data sources store information within their own unique schemas, making reliable and reproducible research challenging. Consequently, researchers have adopted various data models to improve the efficiency of research. Transforming and loading data into these models is a labor-intensive process that can…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…
Authors not listed
The precision of thermodynamic modeling for ionic liquid (IL)–solute systems is fundamentally reliant on the quality of experimental data. However, prevalent databases such as ILThermo frequently exhibit conflicting measurements for the same systems under identical temperature and pressure conditions. These disparities…
Verry Adrian, Intan Rachmita Sari, Hardya Gustada Hikmahrachim
- Abstract: Jakarta is a metropolitan city and among the most dense city in Indonesia. Jakarta has 12 major indicators of standardize health care delivery (Standard Pelayanan Minimum or SPM) derivates from Ministry of Health consists of services related to maternal and neonatal health, school-aged population…
M. TAMER ÖZSU
Work-in-progressThere has been an increasing recognition of the value of data and of data-based decision making. As a consequence, the development of data science as a field of study has intensified in recent years. However, there is no systematic and comprehensive treatment and understanding of data science. This…
Tiqing Liu, Linda Hwang, Stephen K Burley, Carmen I Nitsche + 3 more
BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models, and computational chemistry…
Angela Lopez-del Rio, Sergio Picart, Alexandre Perera-Lluna
In silico analysis of biological activity data has become an essential technique in pharmaceutical development. Specifically, the so-called proteochemometric models aim to share information between targets in machine learning ligand-target activity prediction models. However, bioactivity datasets used in…
Anastasija Nikiforova
Nowadays open data is entering the mainstream - it is free available for every stakeholder and is often used in business decision-making. It is important to be sure data is trustable and error-free as its quality problems can lead to huge losses. The research discusses how (open) data quality could be assessed. It also…