20 papers · ranked by Valyu relevance
Shiyu Zhang, Yajuan Si, John J. Dziak
Background When analyzing randomized controlled trials (RCTs) data, covariate adjustment is often employed to increase the precision of estimated treatment effects. Missing data in covariates, if not handled properly, can result in biased and inefficient estimates. However, the existing literature on handling missing…
Shiyu Zhang, John J. Dziak, Lizbeth Benson, Jamie R. T. Yap + 5 more
The vision of leveraging digital technologies to deliver real-time psychological interventions in everyday settings is realized via just-in-time adaptive interventions (JITAI) - an intervention design that guides the use of rapidly changing information about a person’s internal states and contexts to decide whether and…
Md. Shaddam Hossain Bagmar, Hua Shen
Missing confounders are common in observational studies and present fundamental challenges for causal effect estimation by weakening identification and increasing sensitivity to model misspecification. Within the missing-indicator framework, existing methods rely on a single working model and achieve consistency only…
Yongseok Lee, Walter L. Leite
Researchers using propensity score analysis (PSA) to estimate treatment effects using secondary data may have to handle data that are missing not at random (MNAR). Existing methods for PSA with MNAR data use logistic regression to model the missing data mechanisms, thus requiring manual specification of functional…
Tetiana Gorbach, Tim P Morris, James R Carpenter
While most methods for missing not at random (MNAR) data in regression models address MNAR outcomes assuming fully observed predictors, real-world observational health and longitudinal studies often violate this assumption. This paper compares several approaches for handling MNAR data in linear regression when…
Vaishnavi Nagesh, Lauren Sanders, Sylvain V. Costes, Pinar Avci + 8 more
Missing data is a fundamental challenge in space biology, where high experimental costs, limited sample availability, and tissue allocation constraints produce datasets that are sparse, multimodal, and heterogeneous. We present a systematic four-stage framework for diagnosing, implementing, and validating data…
Yuta Kobayashi, Vincent Jeanselme, Shalmali Joshi
Data collection often reflects human decisions. In healthcare, for instance, a referral for a diagnostic test is influenced by the patient's health, their preferences, available resources, and the practitioner's recommendations. Despite the extensive literature on the informativeness of missingness, its implications on…
Franguridi, Grigory, Moon, Hyungsik Roger
We consider a generalized method of moments framework in which a part of the data vector is missing for some units in a completely unrestricted, potentially endogenous way. In this setup, the parameters of interest are usually only partially identified. We characterize the identified set for such parameters using the…
Xiaopeng Luo, Zexi Tan, Zhuowei Wang
Missing value imputation is a fundamental challenge in machine intelligence, heavily dependent on data completeness. Current imputation methods often handle numerical and categorical attributes independently, overlooking critical interdependencies among heterogeneous features. To address these limitations, we propose a…
Lixing Zhang, Yidong Ouyang, Weifu Li, Shixiang Zhu + 2 more
Missing value imputation is a fundamental task in machine learning, with most existing methods assuming that all missing entries correspond to unobserved regular values. In many real-world datasets, however, missingness may arise from two distinct sources: some entries are meaningfully missing (intrinsically absent and…
Ananthan Nambiar, Carlo Melendez, William Stafford Noble
Multi-omic studies promise a more comprehensive view of biological systems by jointly measuring multiple molecular layers. In practice, however, such datasets are rarely complete: entire molecular modalities may be missing for many samples, and observed modalities often contain substantial feature-level missingness.…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Yixin Shi, Simon Davis, Philip D. Charles, Stephen Taylor + 4 more
Missing values (MVs) remain a significant barrier to reliable proteomics analysis, particularly in single-cell proteomics, where small amounts of starting material and limits in detection drive Missing-Not-At-Random (MNAR) sparsity. Existing imputation methods typically target either Missing-At-Random (MAR) or MNAR…
Mengchun Li, Venkatesh Mallikarjun, Andrew Frey, Emmanuel Ogundimu + 1 more
Mass spectrometry-based label-free proteomics data often suffer from missing values, especially for low-abundance proteins or when a protein is completely absent in one condition. This makes it challenging to estimate fold changes reliably and perform downstream analyses. Traditional imputation methods often show…
Jiaxin Zhang, S. Ghazaleh Dashti, John B. Carlin, Katherine J. Lee + 1 more
Estimating the average causal effect (ACE) using observational data is a key focus in causal inference for which missing data present an important challenge. Multiple imputation (MI) is a widely used method for handling missing data and can yield unbiased estimates when the imputation is compatible with the substantive…
Geoffrey J. McLachlan, Jinran Wu
Semi-supervised learning (SSL) constructs classifiers from datasets in which only a subset of observations is labelled, a situation that naturally arises because obtaining labels often requires expert judgement or costly manual effort. This motivates methods that integrate labelled and unlabelled data within a learning…
Milena Wünsch, Moritz Herrmann, Elisa Noltenius, Mattia Mohr + 2 more
Comparison studies in methodological research are intended to compare methods in an evidence-based manner to help data analysts select a suitable method for their application. To provide trustworthy evidence, they must be carefully designed, implemented, and reported, especially given the many decisions made in…
Authors not listed
While the method of manual inspection reliably produces correct results at the introductory level, it can often appear untidy and unintuitive and lacks teachable depth. This paper presents a novel alternative approach designed to achieve the same outcomes as manual inspection but with enhanced clarity. It serves as…
Prathiksha Ramesh, Maria Fyta
Acute myocardial infarction (AMI) remains one of the leading causes of mortality worldwide, and the following post-effects, such as post-AMI inflammation and tissue repair, involve peripheral blood mononuclear cells playing a critical role. The influence of imputation methods in biological data is assessed with respect…
Alireza Rasoulzadeh, Vignesh Senguttuvan, Warren van Loggerenberg, Richard Border + 1 more
Multiplexed assays of variant effects (MAVEs) measure the functional impact of many protein sequence variants in parallel, potentially covering all possible single amino acid substitutions. Unlike current computational variant effect predictors, MAVEs can reveal the effects of variants under different genetic and…