23 papers · ranked by Valyu relevance
Tlamelo Emmanuel, Thabiso Maupong, Dimane Mpoeleng, Thabo Semong + 2 more
'Banyatsang Mphago' 'Oteng Tabona'] Machine learning has been the corner stone in analysing and extracting information from data and often a problem of missing values is encountered. Missing values occur because of various factors like missing completely at random, missing at random or missing not at random. All these…
Breeshey Roskams-Hieter, Jude Wells, Sara Wade
Missing data persists as a major barrier to data analysis across numerous applications. Recently, deep generative models have been used for imputation of missing data, motivated by their ability to capture highly non-linear and complex relationships in the data. In this work, we investigate the ability of deep models…
Matthew A. Bolt, Samantha MaWhinney, Jack W. Pattee, Kristine M. Erlandson + 2 more
Background Missing data prove troublesome in data analysis; at best they reduce a study’s statistical power and at worst they induce bias in parameter estimates. Multiple imputation via chained equations is a popular technique for dealing with missing data. However, techniques for combining and pooling results from…
Panpan Zhang, Sharon X. Xie
In this paper, we compare the performance of available-case analysis (ACA) and several multiple imputation (MI) approaches for handling missing data problems in longitudinal analysis through estimation bias and relative efficiency. When the missingness of covariates depends on observed responses, ACA produces…
Edoardo Costantini, Kyle M. Lang, Klaas Sijtsma, Tim Reeskens
Multiple Imputation (MI) is one of the most popular approaches to addressing missing values in questionnaires and surveys. MI with multivariate imputation by chained equations (MICE) allows flexible imputation of many types of data. In MICE, for each variable under imputation, the imputer needs to specify which…
James F. Troendle, Aparajita Sur, Eric S. Leifer, Tiffany Powell‐Wiley
We discuss practical aspects of conducting sensitivity analyses for missing data with a repeatedly measured outcome. Our motivation is a SMART trial with a repeatedly measured outcome subject to missingness. We discuss and describe delta-based controlled imputation approaches to conducting sensitivity analyses for such…
Jessica Ryan-Despraz, Amanda Wissler
Missing data is a prevalent problem in bioarchaeological research and imputation could provide a promising solution. This work simulated missingness on a control dataset (481 samples × 41 variables) in order to explore imputation methods for mixed data (qualitative and quantitative data). The tested methods included…
Kin Wai Chan
Multiple imputation (MI) is a technique especially designed for handling missing data in public-use datasets. It allows analysts to perform incompletedata inference straightforwardly by using several already imputed datasets released by the dataset owners. However, the existing MI tests require either a restrictive…
Susana Rafaela Martins, Jacobo de Uña‐Álvarez, María del Carmen Iglesias Pérez
'María del Carmen Iglesias Pérez'] In this work logistic regression when both the response and the predictor variables may be missing is considered. Several existing approaches are reviewed, including complete case analysis, inverse probability weighting, multiple imputation and maximum likelihood. The methods are…
Edoardo Costantini, Kyle M. Lang, Tim Reeskens, Klaas Sijtsma
Including a large number of predictors in the imputation model underlying a multiple imputation (MI) procedure is one of the most challenging tasks imputers face. A variety of high-dimensional MI techniques can help, but there has been limited research on their relative performance. In this study, we investigated a…
Brennan H. Baker, Sheela Sathyanarayana, Adam A. Szpiro, James MacDonald + 1 more
Missing covariate data is a common problem that has not been addressed in observational studies of gene expression. Here we present a multiple imputation (MI) method that accommodates high dimensional transcriptomic data by binning genes, creating separate MI datasets and differential expression models within each bin…
Carlos Traynor, Tarjinder Sahota, Helen Tomkinson, Ignacio Gonzalez-Garcia + 3 more
Missing data is a universal problem in analysing Real-World Evidence (RWE) datasets. In RWE datasets, there is a need to understand which features best correlate with clinical outcomes. In this context, the missing status of several biomarkers may appear as gaps in the dataset that hide meaningful values for analysis.…
Marie Chion, Christine Carapito, Frédèric Bertrand
Motivation: Imputing missing values is common practice in label-free quantitative proteomics. Imputation aims at replacing a missing value with a userdefined one. However, the imputation itself may not be optimally considered downstream of the imputation process, as imputed datasets are often considered as if they had…
Lin Li, Mohammadreza Bayat, Timothy B. Hayes, Wesley K. Thompson + 2 more
This paper addresses the challenges of managing missing values within expansive longitudinal neu-roimaging datasets, using the specific example of data derived from the Adolescent Brain and Cog-nitive Development (ABCD^®^) study. The conventional listwise deletion method, while widely used, is not recommended due to…
Edoardo Costantini, Kyle M. Lang, Klaas Sijtsma, Tim Reeskens
Multiple Imputation (MI) is one of the most popular approaches to addressing missing values in questionnaires and surveys. MI with multivariate imputation by chained equations (MICE) allows flexible imputation of many types of data. In MICE, for each variable under imputation, the imputer needs to specify which…
Yongseok Lee, Walter L. Leite
Researchers using propensity score analysis (PSA) to estimate treatment effects using secondary data may have to handle data that are missing not at random (MNAR). Existing methods for PSA with MNAR data use logistic regression to model the missing data mechanisms, thus requiring manual specification of functional…
Sixia Chen, David Haziza, Victoire Michal
Item nonresponse is a common issue in surveys. Because unadjusted estimators may be biased in the presence of nonresponse, it is common practice to impute the missing values with the objective of reducing the nonresponse bias as much as possible. However, commonly used imputation procedures may lead to unstable…
Kiran H. Kumar, Simone Rubinacci, Sebastian Zöllner
The advent of efficient and accurate imputation for low coverage sequencing offers an unbiased alternative to SNP array imputation, increasing the accuracy of rare variant imputation across all populations. Since imputation accuracy generally increases with larger reference panels and closer ancestry match between…
Barbora Rehák Bučková, Charlotte Fraza, Cecilie Koldbæk Lemvigh, Camilla Bärthel Flaaten + 11 more
Missing data remain a ubiquitous and critical challenge in large-scale clinical studies. Despite advances in imputation, most existing methods fail to address structured missingness, where data are missing according a deterministic pattern and which arise due to systematic patterns introduced by experimental design…
Authors not listed
Solute carrier (SLC) transporters constitute the largest family of membrane transport proteins in humans. They facilitate the movement of ions, neurotransmitters, nutrients, and drugs. Given their critical role in regulating cellular physiology, they are important therapeutic targets for neurological and psychological…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Diba Behnoudfar, Cory Simon, Joshua Schrier
Aqueous, two-phase systems (ATPSs) may form upon mixing two solutions of independently water-soluble compounds. Many separation, purification, and extraction processes rely on ATPSs. Predicting the miscibility of solutions can accelerate and reduce the cost of the discovery of new ATPSs for these applications. Whereas…
Mingze Bai, Jingwen Deng, Chengxin Dai, Julianus Pfeuffer + 1 more
Testing for significant differences in quantities on protein level is a common goal of many LFQ-based mass spectrometry proteomics experiments. Starting from a table of protein and/or peptide quantities from a fixed proteomics quantification software, there exists a multitude of tools and R packages to perform the…