14 papers · ranked by Valyu relevance
Brett K. Beaulieu-Jones, Daniel R. Lavage, John W. Snyder, Jason H. Moore + 2 more
Missing data is a challenge for all studies; however, this is especially true for electronic health record (EHR) based analyses. Failure to appropriately consider missing data can lead to biased results. Here, we provide detailed procedures for when and how to conduct imputation of EHR data. We demonstrate how the…
Daniel W.A. Noble, Shinichi Nakagawa
Ecological and evolutionary research questions are increasingly requiring the integration of research fields along with larger datasets to address fundamental local and global scale problems. Unfortunately, these agendas are often in conflict with limited funding and a need to balance animal welfare concerns. Planned…
Zijian Wang, Jan Hasenauer, Yannik Schälte
Amortized simulation-based neural posterior estimation provides a novel machine learning based approach for solving parameter estimation problems. It has been shown to be computationally efficient and able to handle complex models and data sets. Yet, the available approach cannot handle the in experimental studies…
Mikko Särkkä, Sami Myöhänen, Kaloyan Marinov, Inka Saarinen + 3 more
Modern clinical genetic tests utilize next-generation sequencing (NGS) approaches to comprehensively analyze genetic variants from patients. Out of these millions of variants, clinically relevant variants that match the patient’s phenotype need to be identified accurately within a rapid timeframe that facilitates…
Lin Li, Mohammadreza Bayat, Timothy B. Hayes, Wesley K. Thompson + 2 more
This paper addresses the challenges of managing missing values within expansive longitudinal neu-roimaging datasets, using the specific example of data derived from the Adolescent Brain and Cog-nitive Development (ABCD^®^) study. The conventional listwise deletion method, while widely used, is not recommended due to…
Barbora Rehák Bučková, Charlotte Fraza, Cecilie Koldbæk Lemvigh, Camilla Bärthel Flaaten + 11 more
Missing data remain a ubiquitous and critical challenge in large-scale clinical studies. Despite advances in imputation, most existing methods fail to address structured missingness, where data are missing according a deterministic pattern and which arise due to systematic patterns introduced by experimental design…
Mithilesh Prakash, Jussi Tohka
We introduce a new subtype of ‘Missing Not at Random’ (MNAR) data, where the missingness is correlated with the labels (y) to be predicted, termed (y)-dependent MNAR. We demonstrate that this subtype can significantly bias the estimation of performance metrics in typical machine learning tasks. Unbiased error…
Huiting Ou, Anuradha Surendra, Graeme S.V. McDowell, Emily Hashimoto-Roth + 3 more
Missing values are often unavoidable in modern high-throughput measurements due to various experimental or analytical reasons. Imputation, the process of replacing missing values in a dataset with estimated values, plays an important role in multivariate and machine learning analyses. Three missingness patterns have…
Harvard Wai Hann Hui, Wilson Wen Bin Goh
Statistical analyses in high-dimensional omics data are often hampered by the presence of batch effects (BEs) and missing values (MVs), but the interaction between these two issues is not well-studied nor understood. MVs may manifest as a BE when their proportions differ across batches. These are termed as Batch-Effect…
Tabea Kossen, Michelle Livne, Vince I Madai, Ivana Galinovic + 2 more
Handling missing values is a prevalent challenge in the analysis of clinical data. The rise of data-driven models demands an efficient use of the available data. Methods to impute missing values are thus crucial. Here, we developed a publicly available framework to test different imputation methods and compared their…
Dennis Dimitri Krutkin, Sydney Thomas, Simone Zuffa, Prajit Rajkumar + 3 more
Untargeted metabolomics often produce large datasets with missing values, arising from biological or technical factors, which can undermine statistical analyses and lead to biased biological interpretations. Imputation methods, such as k-Nearest Neighbors (kNN) and Random Forest (RF) regression are commonly used but…
Mustafa Buyukozkan, Elisa Benedetti, Jan Krumsiek
High-dimensional omics datasets frequently contain missing data points, which typically occur due to concentrations below the limit of detection (LOD) of the profiling platform. The presence of such missing values significantly limits downstream statistical analysis and result interpretation. Two common techniques to…
V. Abedi, M.K. Shivakumar, P. Lu, R. Hontecillas + 5 more
Imputation is a key step in Electronic Health Records-mining as it can significantly affect the conclusions derived from the downstream analysis. There are three main categories that explain the missingness in clinical settings–incompleteness, inconsistency, and inaccuracy–and these can capture a variety of situations…
Xing Chen, Na Zhang, Xiaohui Yang, Chunyan Wang + 5 more
In daily life, two common algorithms are used for collecting medical disease data: data integration of medical institutions and questionnaires. However, these statistical methods require collecting data from the entire research area, which consumes a significant amount of manpower and material resources. Additionally…