13 papers · ranked by Valyu relevance
Huiting Ou, Anuradha Surendra, Graeme S.V. McDowell, Emily Hashimoto-Roth + 3 more
Missing values are often unavoidable in modern high-throughput measurements due to various experimental or analytical reasons. Imputation, the process of replacing missing values in a dataset with estimated values, plays an important role in multivariate and machine learning analyses. Three missingness patterns have…
Runmin Wei, Jingye Wang, Erik Jia, Tianlu Chen + 2 more
Left-censored missing values commonly exist in targeted metabolomics datasets and can be considered as missing not at random (MNAR). Improper data processing procedures for missing values will cause adverse impacts on subsequent statistical analyses. However, few imputation methods have been developed and applied to…
Tabea Kossen, Michelle Livne, Vince I Madai, Ivana Galinovic + 2 more
Handling missing values is a prevalent challenge in the analysis of clinical data. The rise of data-driven models demands an efficient use of the available data. Methods to impute missing values are thus crucial. Here, we developed a publicly available framework to test different imputation methods and compared their…
Janine Egert, Bettina Warscheid, Clemens Kreutz
Imputation is a prominent strategy when dealing with missing values (MVs) in proteomics data analysis pipelines. However, the performance of different imputation methods is difficult to assess and varies strongly depending on data characteristics. To overcome this issue, we present the concept of a data-driven…
Mikko Särkkä, Sami Myöhänen, Kaloyan Marinov, Inka Saarinen + 3 more
Modern clinical genetic tests utilize next-generation sequencing (NGS) approaches to comprehensively analyze genetic variants from patients. Out of these millions of variants, clinically relevant variants that match the patient’s phenotype need to be identified accurately within a rapid timeframe that facilitates…
Brett K. Beaulieu-Jones, Daniel R. Lavage, John W. Snyder, Jason H. Moore + 2 more
Missing data is a challenge for all studies; however, this is especially true for electronic health record (EHR) based analyses. Failure to appropriately consider missing data can lead to biased results. Here, we provide detailed procedures for when and how to conduct imputation of EHR data. We demonstrate how the…
Harvard Wai Hann Hui, Wilson Wen Bin Goh
Statistical analyses in high-dimensional omics data are often hampered by the presence of batch effects (BEs) and missing values (MVs), but the interaction between these two issues is not well-studied nor understood. MVs may manifest as a BE when their proportions differ across batches. These are termed as Batch-Effect…
Dennis Dimitri Krutkin, Sydney Thomas, Simone Zuffa, Prajit Rajkumar + 3 more
Untargeted metabolomics often produce large datasets with missing values, arising from biological or technical factors, which can undermine statistical analyses and lead to biased biological interpretations. Imputation methods, such as k-Nearest Neighbors (kNN) and Random Forest (RF) regression are commonly used but…
V. Abedi, M.K. Shivakumar, P. Lu, R. Hontecillas + 5 more
Imputation is a key step in Electronic Health Records-mining as it can significantly affect the conclusions derived from the downstream analysis. There are three main categories that explain the missingness in clinical settings–incompleteness, inconsistency, and inaccuracy–and these can capture a variety of situations…
Mithilesh Prakash, Jussi Tohka
We introduce a new subtype of ‘Missing Not at Random’ (MNAR) data, where the missingness is correlated with the labels (y) to be predicted, termed (y)-dependent MNAR. We demonstrate that this subtype can significantly bias the estimation of performance metrics in typical machine learning tasks. Unbiased error…
Barbora Rehák Bučková, Charlotte Fraza, Cecilie Koldbæk Lemvigh, Camilla Bärthel Flaaten + 11 more
Missing data remain a ubiquitous and critical challenge in large-scale clinical studies. Despite advances in imputation, most existing methods fail to address structured missingness, where data are missing according a deterministic pattern and which arise due to systematic patterns introduced by experimental design…
Lucas Jardim, Luis Mauricio Bini, José Alexandre Felizola Diniz-Filho, Fabricio Villalobos
Given the prevalence of missing data on species’ traits – Raunkiaeran shorfall — and its importance for theoretical and empirical investigations, several methods have been proposed to fill sparse databases. Despite its advantages, imputation of missing data can introduce biases. Here, we evaluate the bias in…
Lin Li, Mohammadreza Bayat, Timothy B. Hayes, Wesley K. Thompson + 2 more
This paper addresses the challenges of managing missing values within expansive longitudinal neu-roimaging datasets, using the specific example of data derived from the Adolescent Brain and Cog-nitive Development (ABCD^®^) study. The conventional listwise deletion method, while widely used, is not recommended due to…