23 papers · ranked by Valyu relevance
Jicong Fan
Missing data represents a fundamental and pervasive challenge in modern data science, significantly impeding analytical capabilities and decision-making processes across an exceptionally broad spectrum of disciplines including healthcare, bioinformatics, social science, e-commerce, and industrial monitoring systems.…
Pegah Golchian, Jan Kapar, David S. Watson, Marvin N. Wright
Handling missing values is a common challenge in biostatistical analyses, typically addressed by imputation methods. We propose a novel, fast, and easy-to-use imputation method called missing value imputation with adversarial random forests (MissARF), based on generative machine learning, that provides both single and…
Hugo Morvan, Jonas Agholme, Bjorn Eliasson, Katarina Olofsson + 3 more
Missing data is a prevalent issue in many applications, including large medical registries such as the Swedish Healthcare Quality Registries, potentially leading to biased or inefficient analyses if not handled properly. Multiple Imputation by Chained Equations (MICE) is a popular and versatile method for handling…
George Sun, Yi-Hui Zhou
In this study, we introduce a sophisticated generative conditional strategy designed to impute missing values within datasets, an area of considerable importance in statistical analysis. Specifically, we initially elucidate the theoretical underpinnings of the Generative Conditional Missing Imputation Networks (GCMI)…
Niki Z. Petrakos, Erica E. M. Moodie, Nicolas Savy
The current literature regarding generation of complex, realistic synthetic tabular data, particularly for randomized controlled trials (RCTs), often ignores missing data. However, missing data are common in RCT data and often are not Missing Completely At Random. We bridge the gap of determining how best to generate…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Lodato, Ivano, Iyer, Aditya V. + 2 more
We introduce a method for evaluating interventional queries and Average Treatment Effects (ATEs) in the presence of generalized incomplete contingency tables (GICTs), contingency tables containing a full row of random (sampling) zeros, rendering some conditional probabilities undefined. Rather than discarding such…
Yongseok Lee, Walter L. Leite
Researchers using propensity score analysis (PSA) to estimate treatment effects using secondary data may have to handle data that are missing not at random (MNAR). Existing methods for PSA with MNAR data use logistic regression to model the missing data mechanisms, thus requiring manual specification of functional…
Aya El Mir, Eric Bezerra de Sousa, Ignacio Mesina-Estarrón, Leo Anthony Celi + 8 more
Missing, inaccurate, or poorly documented data in healthcare is often treated as a technical problem to be statistically resolved via imputation, deletion, or modeling assumptions about randomness. However, such inaccuracies relate to far more complex socioeconomic and geopolitical issues, rather than “errors of data…
Asmaa Ahmad, Eric J. Rose, Michael S. Roy, Edward Valachovic + 1 more
Missing data in periodic time series can bias inference when temporal dependence is not preserved during imputation. We propose a framework that integrates the Variable Bandpass Periodic Block Bootstrap (VBPBB) with multiple imputation using Amelia II by incorporating statistically significant periodic components as…
Ananthan Nambiar, Carlo Melendez, William Stafford Noble
Multi-omic studies promise a more comprehensive view of biological systems by jointly measuring multiple molecular layers. In practice, however, such datasets are rarely complete: entire molecular modalities may be missing for many samples, and observed modalities often contain substantial feature-level missingness.…
Paul Madley-Dowd, Rachael A Hughes, Maya B Mathur, Jon Heron + 1 more
Missing data are a pervasive problem in epidemiology, with multiple imputation (MI) a commonly used analysis method. MI is valid when data are missing at random (MAR). However, definitions of MAR with multiple incomplete variables are not easily interpretable and descriptions of graphical model-based conditions are not…
Mengchun Li, Venkatesh Mallikarjun, Andrew Frey, Emmanuel Ogundimu + 1 more
Mass spectrometry-based label-free proteomics data often suffer from missing values, especially for low-abundance proteins or when a protein is completely absent in one condition. This makes it challenging to estimate fold changes reliably and perform downstream analyses. Traditional imputation methods often show…
Harlan Campbell, Tim P. Morris, Paul Gustafson
Derived variables are variables that are constructed from one or more source variables through established mathematical operations or algorithms. For example, body mass index (BMI) is a derived variable constructed from two source variables: weight and height. When using a derived variable as the outcome in a…
Yixin Shi, Simon Davis, Philip D. Charles, Stephen Taylor + 4 more
Missing values (MVs) remain a significant barrier to reliable proteomics analysis, particularly in single-cell proteomics, where small amounts of starting material and limits in detection drive Missing-Not-At-Random (MNAR) sparsity. Existing imputation methods typically target either Missing-At-Random (MAR) or MNAR…
Vaishnavi Nagesh, Lauren Sanders, Sylvain V. Costes, Pinar Avci + 8 more
Missing data is a fundamental challenge in space biology, where high experimental costs, limited sample availability, and tissue allocation constraints produce datasets that are sparse, multimodal, and heterogeneous. We present a systematic four-stage framework for diagnosing, implementing, and validating data…
Adrienne Kline, Yuan Luo
Most datasets suffer from partial or complete missing values, which has downstream limitations on the available models on which to test the data and on any statistical inferences that can be made from the data. Several imputation techniques have been designed to replace missing data with stand in values. The various…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Jörn Lötsch, Alfred Ultsch
Missing value imputation is a routine step in biomedical data analysis, yet techniques are often not tailored to specific datasets. We propose a systematic framework for selecting imputation methods customized for the unique characteristics of cross-sectional numerical data, with a focus on pain-related biomedical…
Authors not listed
Predicting solution conformation and aggregation of conjugated polymers remains a bottleneck for translating solution processing into controlled film microstructure and for closing the loop in self-driving laboratories. We construct a cleaned, machine-readable dataset of 256 entries that links polymer size…
T. Quinn Smith, Amatur Rahman, Zachary A. Szpiech
We introduce Empirical Genotype Generalizer for Samples (EGGS) which accepts empirical genotypes with missing data and replicates the distribution of missing genotypes along the empirical segment in other replicates. The empirical segment must have a number of sites less than the replicate. In addition, EGGS can remove…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Authors not listed
The analysis of metabolic profiles using high resolution mass spectrometry (MS) data gives deep insights into the biological processes. In metabolomics, MS generates a large number of features that represent metabolites. However, identifying specific metabolites from these features can be challenging. One of the major…