13 papers · ranked by Valyu relevance
Mike Van Ness, Tomas M. Bosschieter, Roberto Halpin-Gregorio, Madeleine Udell
'Madeleine Udell'] Missing data is common in applied data science, particularly for tabular data sets found in healthcare, social sciences, and natural sciences. Most supervised learning methods only work on complete data, thus requiring preprocessing such as missing value imputation to work on incomplete data sets.…
Shiyu Zhang, Yajuan Si, John J. Dziak
Background When analyzing randomized controlled trials (RCTs) data, covariate adjustment is often employed to increase the precision of estimated treatment effects. Missing data in covariates, if not handled properly, can result in biased and inefficient estimates. However, the existing literature on handling missing…
Molly Ehrig, Garrett S Bullock, Xiaoyan Iris Leng, Nicholas M Pajewski + 2 more
'Nicholas M Pajewski' 'Jaime Lynn Speiser' 'Christian Lovis'] Title: Abstract Background Missing data in electronic health records are highly prevalent and result in analytical concerns such as heterogeneous sources of bias and loss of statistical power. One simple analytic method for addressing missing or unknown…
Mutamba T. Kayembe, Shahab Jolani, Frans E. S. Tan, Gerard J. P. van Breukelen
'Gerard J. P. van Breukelen'] Title: Summary In this article, we first review the literature on dealing with missing values on a covariate in randomized studies and summarize what has been done and what is lacking to date. We then investigate the situation with a continuous outcome and a missing binary covariate in…
Oliver Urs Lenz, Daniel Peralta, Chris Cornelis
Imputation allows datasets to be used with algorithms that cannot handle missing values by themselves. However, missing values may in principle contribute useful information that is lost through imputation. The missing-indicator approach can be used to preserve this information. There are several theoretical…
Matthew Sperrin, Glen P. Martin
Background Within routinely collected health data, missing data for an individual might provide useful information in itself. This occurs, for example, in the case of electronic health records, where the presence or absence of data is informative. While the naive use of missing indicators to try to exploit such…
Gift Khangamwa, Terence L. van Zyl, C. J. van Alten
Missing data is a common concern in health datasets, and its impact on good decision-making processes is well documented. Our study's contribution is a methodology for tackling missing data problems using a combination of synthetic dataset generation, missing data imputation and deep learning methods to resolve missing…
Mingyang Song, Xin Zhou, Mathew J. Pazaris, Donna Spiegelman
1 Departments of Epidemiology and Nutrition, Harvard T.H. Chan School of Public Health, Boston, MA, USA. 2 Clinical and Translational Epidemiology Unit, Mongan Institute, Massachusetts General Hospital, Boston, MA, USA. 3 Division of Gastroenterology, Massachusetts General Hospital and Harvard Medical School, Boston…
Anqi Zhao, Peng Ding
Complete randomization allows for consistent estimation of the average treatment effect based on the difference in means of the outcomes without strong modeling assumptions on the outcome-generating process. Appropriate use of the pretreatment covariates can further improve the estimation efficiency. However…
Morten Wærsted, Taran Svenssen Børnick, Jos W. R. Twisk, Kaj Bo Veiersted
'Kaj Bo Veiersted'] Objective Missing data in longitudinal studies may constitute a source of bias. We suggest three simple missing data indicators for the initial phase of getting an overview of the missingness pattern in a dataset with a high number of follow-ups. Possible use of the indicators is exemplified in two…
Shan Gao, Elena Albu, Pieter Stijnen, Frank Rademakers + 5 more
- 1 Department of Development and Regeneration, KU Leuven, Leuven, Belgium - 2 Management Information Reporting Department, University Hospitals Leuven, Leuven, Belgium - 3 Faculty of Medicine, KU Leuven, Leuven, Belgium - 4 Department of Infection Control and Prevention, University Hospitals Leuven, Leuven, Belgium -…
Ayman Omar Baniamer, Henri Tilga
Statistical models are essential tools in data analysis. However, missing data plays a pivotal role in impacting the assumptions and effectiveness of statistical models, especially when there is a significant amount of missing data. This study addresses one of the core assumptions supporting many statistical models…
M. Templ, Markus Ulmer
Many imputation methods have been developed over the years and tested mostly under ideal settings. Surprisingly, there is no detailed research on how imputation methods perform when the idealized assumptions about the distribution of data and/or model assumptions are partly not fulfilled. This research looks into the…