16 papers · ranked by Valyu relevance
Gift Khangamwa, Terence L. van Zyl, C. J. van Alten
Missing data is a common concern in health datasets, and its impact on good decision-making processes is well documented. Our study's contribution is a methodology for tackling missing data problems using a combination of synthetic dataset generation, missing data imputation and deep learning methods to resolve missing…
Youran Zhou, Mohamed Reda Bouadjenek, Sunil Aryal
Missing data is a pervasive challenge spanning diverse data types, including tabular, sensor data, time-series, images and so on. Its origins are multifaceted, resulting in various missing mechanisms. Prior research in this field has predominantly revolved around the assumption of the Missing Completely At Random…
Jahan C. Penny‐Dimri, Christoph Bergmeir, Julian A. Smith
Most practical data science problems encounter missing data. A wide variety of solutions exist, each with strengths and weaknesses that depend upon the missingness-generating process. Here we develop a theoretical framework for training and inference using only observed variables enabling modeling of incomplete…
Youran Zhou, Sunil Aryal, Mohamed Reda Bouadjenek
Missing data poses a significant challenge in data science, affecting decision-making processes and outcomes. Understanding what missing data is, how it occurs, and why it is crucial to handle it appropriately is paramount when working with real-world data, especially in tabular data, one of the most commonly used data…
Azad, Fatemeh, Bosnic Zoran, Kukar + 1 more
—Missing data represents a fundamental challenge in machine learning applications, often reducing model performance and reliability. This problem is particularly acute in fields like bioinformatics and clinical machine learning, where datasets are frequently incomplete due to the nature of both data generation and data…
Sara Johansson Fernstad, Sarah Alsufyani, Silvia Del Din, Alison Yarnall + 1 more
'Alison Yarnall' 'Lynn Rochester'] This paper contributes a set of quality metrics for identification and visual analysis of structured missingness in high-dimensional data. Missing values in data are a frequent challenge in most data generating domains and may cause a range of analysis issues. Structural missingness…
Neslihan Süzen, Evgeny M. Mirkes, Damian Roland, Jeremy Levesley + 2 more
'Alexander N. Gorban' 'Tim Coats'] Abstract— Electronic patient records (EPRs) produce a wealth of data but contain significant missing information. Understanding and handling this missing data is an important part of clinical data analysis and if left unaddressed could result in bias in analysis and distortion in…
Seongmin Kim, Jaewon Oh, Hye-Young Ko, Jeongmin Park + 1 more
Missing data is a common issue in various fields such as medicine, social sciences, and natural sciences, and it poses significant challenges for accurate statistical analysis. Although numerous imputation methods have been proposed to address this issue, many of them fail to adequately capture the complex dependency…
Muhammad Ishaq, Laila iftikhar, Majid Khan, Asfandyar Khan + 1 more
'Arshad Khan'] This study explored the use of machine learning algorithms for predicting and imputing missing values in categorical datasets. We focused on ensemble models that use the error correction output codes (ECOC) framework, including SVM-based and KNN-based ensemble models, as well as an ensemble classifier…
Hugo Morvan, Jonas Agholme, Bjorn Eliasson, Katarina Olofsson + 3 more
Missing data is a prevalent issue in many applications, including large medical registries such as the Swedish Healthcare Quality Registries, potentially leading to biased or inefficient analyses if not handled properly. Multiple Imputation by Chained Equations (MICE) is a popular and versatile method for handling…
Sarah Alsufyani, Matthew Forshaw, Sara Johansson Fernstad
Missing data, the data value that is not recorded for a variable, occurs in almost all statistical analyses and may be caused by many reasons, such as lack of collection or a lack of documentation. Researchers need to adequately deal with this issue to provide a valid analysis. The visualization of missing values plays…
Lodato, Ivano, Iyer, Aditya V. + 2 more
We introduce a method for evaluating interventional queries and Average Treatment Effects (ATEs) in the presence of generalized incomplete contingency tables (GICTs), contingency tables containing a full row of random (sampling) zeros, rendering some conditional probabilities undefined. Rather than discarding such…
Dandan Tang, Xin Tong
Research: Traditional and Machine Learning Approaches Authors: ['Dandan Tang' 'Xin Tong'] Missing Not at Random (MNAR) and nonnormal data are challenging to handle. Traditional missing data analytical techniques such as full information maximum likelihood estimation (FIML) may fail with nonnormal data as they are built…
Niki Z. Petrakos, Erica E. M. Moodie, Nicolas Savy
The current literature regarding generation of complex, realistic synthetic tabular data, particularly for randomized controlled trials (RCTs), often ignores missing data. However, missing data are common in RCT data and often are not Missing Completely At Random. We bridge the gap of determining how best to generate…
Tu T. Do, Mai-Anh Vu, Hoang Thien Ly, Thu Nguyen + 4 more
'Michael A. Riegler' 'Pål Halvorsen' 'Binh T. Nguyen'] Abstract—Monotone missing data is a common problem in data analysis. However, imputation combined with dimensionality reduction can be computationally expensive, especially with the increasing size of datasets. To address this issue, we propose a Blockwise…
Daniel Zhang
Time series data are observations collected over time intervals. Successful analysis of time series data captures patterns such as trends, cyclicity and irregularity, which are crucial for decision making in research, business, and governance. However, missing values in time series data occur often and present…