23 papers · ranked by Valyu relevance
Mike Van Ness, Tomas M. Bosschieter, Roberto Halpin-Gregorio, Madeleine Udell
'Madeleine Udell'] Missing data is common in applied data science, particularly for tabular data sets found in healthcare, social sciences, and natural sciences. Most supervised learning methods only work on complete data, thus requiring preprocessing such as missing value imputation to work on incomplete data sets.…
Shiyu Zhang, Yajuan Si, John J. Dziak
Background When analyzing randomized controlled trials (RCTs) data, covariate adjustment is often employed to increase the precision of estimated treatment effects. Missing data in covariates, if not handled properly, can result in biased and inefficient estimates. However, the existing literature on handling missing…
Molly Ehrig, Garrett S Bullock, Xiaoyan Iris Leng, Nicholas M Pajewski + 2 more
'Nicholas M Pajewski' 'Jaime Lynn Speiser' 'Christian Lovis'] Title: Abstract Background Missing data in electronic health records are highly prevalent and result in analytical concerns such as heterogeneous sources of bias and loss of statistical power. One simple analytic method for addressing missing or unknown…
Mutamba T. Kayembe, Shahab Jolani, Frans E. S. Tan, Gerard J. P. van Breukelen
'Gerard J. P. van Breukelen'] Title: Summary In this article, we first review the literature on dealing with missing values on a covariate in randomized studies and summarize what has been done and what is lacking to date. We then investigate the situation with a continuous outcome and a missing binary covariate in…
Oliver Urs Lenz, Daniel Peralta, Chris Cornelis
Imputation allows datasets to be used with algorithms that cannot handle missing values by themselves. However, missing values may in principle contribute useful information that is lost through imputation. The missing-indicator approach can be used to preserve this information. There are several theoretical…
Matthew Sperrin, Glen P. Martin
Background Within routinely collected health data, missing data for an individual might provide useful information in itself. This occurs, for example, in the case of electronic health records, where the presence or absence of data is informative. While the naive use of missing indicators to try to exploit such…
Gift Khangamwa, Terence L. van Zyl, C. J. van Alten
Missing data is a common concern in health datasets, and its impact on good decision-making processes is well documented. Our study's contribution is a methodology for tackling missing data problems using a combination of synthetic dataset generation, missing data imputation and deep learning methods to resolve missing…
Mingyang Song, Xin Zhou, Mathew J. Pazaris, Donna Spiegelman
1 Departments of Epidemiology and Nutrition, Harvard T.H. Chan School of Public Health, Boston, MA, USA. 2 Clinical and Translational Epidemiology Unit, Mongan Institute, Massachusetts General Hospital, Boston, MA, USA. 3 Division of Gastroenterology, Massachusetts General Hospital and Harvard Medical School, Boston…
Anqi Zhao, Peng Ding
Complete randomization allows for consistent estimation of the average treatment effect based on the difference in means of the outcomes without strong modeling assumptions on the outcome-generating process. Appropriate use of the pretreatment covariates can further improve the estimation efficiency. However…
Morten Wærsted, Taran Svenssen Børnick, Jos W. R. Twisk, Kaj Bo Veiersted
'Kaj Bo Veiersted'] Objective Missing data in longitudinal studies may constitute a source of bias. We suggest three simple missing data indicators for the initial phase of getting an overview of the missingness pattern in a dataset with a high number of follow-ups. Possible use of the indicators is exemplified in two…
Shan Gao, Elena Albu, Pieter Stijnen, Frank Rademakers + 5 more
- 1 Department of Development and Regeneration, KU Leuven, Leuven, Belgium - 2 Management Information Reporting Department, University Hospitals Leuven, Leuven, Belgium - 3 Faculty of Medicine, KU Leuven, Leuven, Belgium - 4 Department of Infection Control and Prevention, University Hospitals Leuven, Leuven, Belgium -…
Mikko Särkkä, Sami Myöhänen, Kaloyan Marinov, Inka Saarinen + 3 more
Modern clinical genetic tests utilize next-generation sequencing (NGS) approaches to comprehensively analyze genetic variants from patients. Out of these millions of variants, clinically relevant variants that match the patient’s phenotype need to be identified accurately within a rapid timeframe that facilitates…
Mustafa Buyukozkan, Elisa Benedetti, Jan Krumsiek
High-dimensional omics datasets frequently contain missing data points, which typically occur due to concentrations below the limit of detection (LOD) of the profiling platform. The presence of such missing values significantly limits downstream statistical analysis and result interpretation. Two common techniques to…
Huiting Ou, Anuradha Surendra, Graeme S.V. McDowell, Emily Hashimoto-Roth + 3 more
Missing values are often unavoidable in modern high-throughput measurements due to various experimental or analytical reasons. Imputation, the process of replacing missing values in a dataset with estimated values, plays an important role in multivariate and machine learning analyses. Three missingness patterns have…
Diba Behnoudfar, Cory Simon, Joshua Schrier
Aqueous, two-phase systems (ATPSs) may form upon mixing two solutions of independently water-soluble compounds. Many separation, purification, and extraction processes rely on ATPSs. Predicting the miscibility of solutions can accelerate and reduce the cost of the discovery of new ATPSs for these applications. Whereas…
Kruttika Dabke, Simion Kreimer, Michelle R. Jones, Sarah J. Parker
Missing values in proteomic data sets have real consequences on downstream data analysis and reproducibility. Although several imputation methods exist to handle missing values, there is no single imputation method that is best suited for a diverse range of data sets and no clear strategy exists for evaluating…
Runmin Wei, Jingye Wang, Erik Jia, Tianlu Chen + 2 more
Left-censored missing values commonly exist in targeted metabolomics datasets and can be considered as missing not at random (MNAR). Improper data processing procedures for missing values will cause adverse impacts on subsequent statistical analyses. However, few imputation methods have been developed and applied to…
Ayman Omar Baniamer, Henri Tilga
Statistical models are essential tools in data analysis. However, missing data plays a pivotal role in impacting the assumptions and effectiveness of statistical models, especially when there is a significant amount of missing data. This study addresses one of the core assumptions supporting many statistical models…
Xuhua Xia
Missing data are frequently encountered in molecular phylogenetics and need to be imputed. For a distance matrix with missing distances, the least- squares approach is often used for imputing the missing values. Here I develop a method, similar to the expectation-maximization algorithm, to impute multiple missing…
M. Templ, Markus Ulmer
Many imputation methods have been developed over the years and tested mostly under ideal settings. Surprisingly, there is no detailed research on how imputation methods perform when the idealized assumptions about the distribution of data and/or model assumptions are partly not fulfilled. This research looks into the…
Julie Macleod, Stephen Thomas
‘Hidden’ catalysis plagues the development and understanding of all catalytic processes. Hidden acid catalysis and catalysis by trace metal contamination being two widely recognised examples. Since 2010, over 600 new catalysed hydroboration protocols have been reported despite the prevalence of hidden borane catalysis…
Tieu-Long Phan, Klaus Weinbauer, Thomas Gärtner, Daniel Merkle + 3 more
Purpose: Reaction databases are a key resource for a wide variety of applications in computational chemistry and biochemistry, including Computer-aided Synthesis Planning (CASP) and the large-scale analysis of metabolic networks. The full potential of these resources can only be realized if datasets are accurate and…
Justin Eilertsen, Wylie Stroberg, Santiago Schnell
The determination of a substrate or enzyme activity by coupling of one enzymatic reaction with another easily detectable (indicator) reaction is a common practice in the biochemical sciences. Usually, the kinetics of enzyme reactions is simplified with singular perturbation analysis to derive rate or time course…