15 papers · ranked by Valyu relevance
Jicong Fan
Missing data represents a fundamental and pervasive challenge in modern data science, significantly impeding analytical capabilities and decision-making processes across an exceptionally broad spectrum of disciplines including healthcare, bioinformatics, social science, e-commerce, and industrial monitoring systems.…
Hugo Morvan, Jonas Agholme, Bjorn Eliasson, Katarina Olofsson + 3 more
Missing data is a prevalent issue in many applications, including large medical registries such as the Swedish Healthcare Quality Registries, potentially leading to biased or inefficient analyses if not handled properly. Multiple Imputation by Chained Equations (MICE) is a popular and versatile method for handling…
Lukas Klein, Gunter Grieser, Carl-Ludwig Fischer-Fröhlich, Axel Rahmel + 3 more
This study presents an Initial Data Analysis (IDA) of the German Transplantation Registry (TxReg) data for a better data understanding and to inform future data analyses. The IDA is focusing on data on first-time kidney-only transplantations in adult recipients from deceased donors between 2006 and 2016 and refers to…
George Sun, Yi-Hui Zhou
In this study, we introduce a sophisticated generative conditional strategy designed to impute missing values within datasets, an area of considerable importance in statistical analysis. Specifically, we initially elucidate the theoretical underpinnings of the Generative Conditional Missing Imputation Networks (GCMI)…
Niki Z. Petrakos, Erica E. M. Moodie, Nicolas Savy
The current literature regarding generation of complex, realistic synthetic tabular data, particularly for randomized controlled trials (RCTs), often ignores missing data. However, missing data are common in RCT data and often are not Missing Completely At Random. We bridge the gap of determining how best to generate…
Abel Villa, Llorenç Badiella
The analysis of randomized trials is often complicated by the occurrence of intercurrent events and missing values. Even though there are different strategies to address missing values it is still common to require missing values imputation. In the present article we explore the estimation of treatment effects in RCTs…
Durga Keshav, GVD Praneeth, Chetan Kumar Patruni, Vivek Yelleti + 1 more
The missing data problem is one of the important issues to address for achieving data quality. While imputation-based methods are designed to achieve data completeness, their efficacy is observed to be diminishing as and when there is increasing in the missingness percentage. Further, extant approaches often struggle…
Lodato, Ivano, Iyer, Aditya V. + 2 more
We introduce a method for evaluating interventional queries and Average Treatment Effects (ATEs) in the presence of generalized incomplete contingency tables (GICTs), contingency tables containing a full row of random (sampling) zeros, rendering some conditional probabilities undefined. Rather than discarding such…
Wenlu Tang, Hongni Wang, Xingcai Zhou, Bei Jiang + 1 more
We develope a novel approach to tackle the common but challenging problem of conformal inference for missing data in machine learning, focusing on 'Missing at Random' (MAR) data. Our method, inspired by the innovative Multiple Robust (MR) framework, introduces a general approach to construct prediction intervals, which…
Yanjiao Yang, Yikun Zhang, Xinwei Shen, Yen-Chi Chen
We propose Emputation, a deep generative framework for learning imputation models. Emputation targets the extrapolation distribution of missing variables given observed variables, and training is guided by specific missingness assumptions that guarantee identification of the target distribution. The training objective…
Xin Guan
The classical $k$-means clustering, based on distances computed from all data features, cannot be directly applied to incomplete data with missing values. A natural extension of $k$-means to missing data is to involve only the observed positions in clustering, which is equivalent to imputing missing values by…
Santu Mondal, Chayan Maitra, Rajat K. De
In real-world machine learning applications, incomplete observations create a fundamental challenge. Researchers have come up with several ideas to address this crucial problem. However, current models still face challenges in balancing scalability and structural consistency. This study proposes a feature imputation…
Lixing Zhang, Yidong Ouyang, Weifu Li, Shixiang Zhu + 2 more
Missing value imputation is a fundamental task in machine learning, with most existing methods assuming that all missing entries correspond to unobserved regular values. In many real-world datasets, however, missingness may arise from two distinct sources: some entries are meaningfully missing (intrinsically absent and…
Zeng, Zhenghao, Arbour, David + 8 more
Human annotations play a crucial role in evaluating the performance of GenAI models. Two common challenges in practice, however, are missing annotations (the response variable of interest) and cluster dependence among human-AI interactions (e.g., questions asked by the same user may be highly correlated). Reliable…
Ben Swallow, Lars Brestrich, Victor Velasco-Pardo
Missing data and confounding are common in real-world statistical applications, yet few studies have examined how imputation methods perform under time-varying confounding in binary variables, or how missingness mechanism, missing rate, missingness location and sample size jointly affect performance and the underlying…