16 papers · ranked by Valyu relevance
Dawei Liu, Hanne Oberman, Johanna Muñoz, Jeroen Hoogland + 1 more
'Thomas P. A. Debray'] This is a preprint of the following chapter: Liu D, Oberman HI, Muñoz J, Hoogland J, Debray TPA, "Quality control, data cleaning, imputation", published in "Clinical applications of artificial intelligence in real-world data", edited by [editor of the book], [year of publication], [publisher (as…
Youran Zhou, Mohamed Reda Bouadjenek, Sunil Aryal
Missing data is a pervasive challenge spanning diverse data types, including tabular, sensor data, time-series, images and so on. Its origins are multifaceted, resulting in various missing mechanisms. Prior research in this field has predominantly revolved around the assumption of the Missing Completely At Random…
Youran Zhou, Sunil Aryal, Mohamed Reda Bouadjenek
Missing data poses a significant challenge in data science, affecting decision-making processes and outcomes. Understanding what missing data is, how it occurs, and why it is crucial to handle it appropriately is paramount when working with real-world data, especially in tabular data, one of the most commonly used data…
Azad, Fatemeh, Bosnic Zoran, Kukar + 1 more
—Missing data represents a fundamental challenge in machine learning applications, often reducing model performance and reliability. This problem is particularly acute in fields like bioinformatics and clinical machine learning, where datasets are frequently incomplete due to the nature of both data generation and data…
Sara Johansson Fernstad, Sarah Alsufyani, Silvia Del Din, Alison Yarnall + 1 more
'Alison Yarnall' 'Lynn Rochester'] This paper contributes a set of quality metrics for identification and visual analysis of structured missingness in high-dimensional data. Missing values in data are a frequent challenge in most data generating domains and may cause a range of analysis issues. Structural missingness…
Nicholas Tierney, Dianne H Cook
Despite the large body of research on missing value distributions and imputation, there is comparatively little literature with a focus on how to make it easy to handle, explore, and impute missing values in data. This paper addresses this gap. The new methodology builds upon tidy data principles, with the goal of…
Neslihan Süzen, Evgeny M. Mirkes, Damian Roland, Jeremy Levesley + 2 more
'Alexander N. Gorban' 'Tim Coats'] Abstract— Electronic patient records (EPRs) produce a wealth of data but contain significant missing information. Understanding and handling this missing data is an important part of clinical data analysis and if left unaddressed could result in bias in analysis and distortion in…
Sara Johansson Fernstad, Jimmy Johansson
—This paper contributes a novel visualization method, Missingness Glyph, for analysis and exploration of missing values in data. Missing values are a common challenge in most data generating domains and may cause a range of analysis issues. Missingness in data may indicate potential problems in data collection and…
Seongmin Kim, Jeunghun Oh, Hungkuk Ko, Jeongmin Park + 1 more
Missing data is a common issue in various fields such as medicine, social sciences, and natural sciences, and it poses significant challenges for accurate statistical analysis. Although numerous imputation methods have been proposed to address this issue, many of them fail to adequately capture the complex dependency…
Gift Khangamwa, Terence L. van Zyl, C. J. van Alten
Missing data is a common concern in health datasets, and its impact on good decision-making processes is well documented. Our study's contribution is a methodology for tackling missing data problems using a combination of synthetic dataset generation, missing data imputation and deep learning methods to resolve missing…
Jahan C. Penny‐Dimri, Christoph Bergmeir, Julian A. Smith
Most practical data science problems encounter missing data. A wide variety of solutions exist, each with strengths and weaknesses that depend upon the missingness-generating process. Here we develop a theoretical framework for training and inference using only observed variables enabling modeling of incomplete…
Alexander Franks, Edoardo M. Airoldi, Donald B. Rubin
∗ Alexander M. Franks is a Moore-Sloan Data Science Fellow at the University of Washington and a graduate of the Department of Statistics at Harvard University (afranks@post.harvard.edu) Edoardo M. Airoldi is an Associate Professor of Statistics at Harvard University (airoldi@fas.harvard.edu). Donald B. Rubin is the…
Maoyuan Sun, Yue Ma, Yuanxin Wang, Tianyi Li + 3 more
'Ping‐Shou Zhong'] Data-driven decision making has been a common task in today's big data era, from simple choices such as finding a fast way to drive home, to complex decisions on medical treatment. It is often supported by visual analytics. For various reasons (e.g., system failure, interrupted network, intentional…
Collins Achepsah Leke, Tshilidzi Marwala, Satyakama Paul
In the last couple of decades, there has been major advancements in the domain of missing data imputation. The techniques in the domain include amongst others: Expectation Maximization, Neural Networks with Evolutionary Algorithms or optimization techniques and K-Nearest Neighbor approaches to solve the problem. The…
Seema Sangari, Herman E. Ray
Missing data is a common problem which has consistently plagued statisticians and applied analytical researchers. While replacement methods like mean-based or hot deck imputation have been well researched, emerging imputation techniques enabled through improved computational resources have had limited formal…
Davi E. N. Frossard, Igor O. Nunes, Renato A. Krohling
Techniques such as clusterization, neural networks and decision making usually rely on algorithms that are not well suited to deal with missing values. However, real world data frequently contains such cases. The simplest solution is to either substitute them by a best guess value or completely disregard the missing…