22 papers · ranked by Valyu relevance
Tlamelo Emmanuel, Thabiso Maupong, Dimane Mpoeleng, Thabo Semong + 2 more
'Banyatsang Mphago' 'Oteng Tabona'] Machine learning has been the corner stone in analysing and extracting information from data and often a problem of missing values is encountered. Missing values occur because of various factors like missing completely at random, missing at random or missing not at random. All these…
Luis Alejandro Masmela-Caita, Thais P. Galletti, Marcos O. Prates
Missing data theory deals with the statistical methods in the occurrence of missing data. Missing data occurs when some values are not stored or observed for variables of interest. However, most of the statistical theory assumes that data is fully observed. An alternative to deal with incomplete databases is to fill in…
Pegah Golchian, Jan Kapar, David S. Watson, Marvin N. Wright
Handling missing values is a common challenge in biostatistical analyses, typically addressed by imputation methods. We propose a novel, fast, and easy-to-use imputation method called missing value imputation with adversarial random forests (MissARF), based on generative machine learning, that provides both single and…
Charles F. Manski
Incomplete observability of data generates an identification problem. There is no panacea for missing data. What one can learn about a population parameter depends on the assumptions one finds credible to maintain. The credibility of assumptions varies with the empirical setting. No specific assumptions can provide a…
Huiting Ou, Anuradha Surendra, Graeme S.V. McDowell, Emily Hashimoto-Roth + 3 more
Missing values are often unavoidable in modern high-throughput measurements due to various experimental or analytical reasons. Imputation, the process of replacing missing values in a dataset with estimated values, plays an important role in multivariate and machine learning analyses. Three missingness patterns have…
Huiting Ou, Anuradha Surendra, Graeme S V McDowell, Emily Hashimoto-Roth + 4 more
Missing data are a major problem for multivariate, machine learning (ML) and network analyses. For example, in large lipidomic or metabolic datasets, measurements for some analytes may not be available in every sample due to routine technical variability, low abundance, ion suppression from co-eluting analytes…
Robert Thiesmeier, Matteo Bottai, Nicola Orsini
Missing data is a common challenge across scientific disciplines. Current imputation methods require the availability of individual data to impute missing values. Often, however, missingness requires using external data for the imputation. In this paper, we introduce a new Stata command, mi impute from, designed to…
Mikko Särkkä, Sami Myöhänen, Kaloyan Marinov, Inka Saarinen + 3 more
Modern clinical genetic tests utilize next-generation sequencing (NGS) approaches to comprehensively analyze genetic variants from patients. Out of these millions of variants, clinically relevant variants that match the patient’s phenotype need to be identified accurately within a rapid timeframe that facilitates…
Ayman Omar Baniamer, Henri Tilga
Statistical models are essential tools in data analysis. However, missing data plays a pivotal role in impacting the assumptions and effectiveness of statistical models, especially when there is a significant amount of missing data. This study addresses one of the core assumptions supporting many statistical models…
Youran Zhou, Sunil Aryal, Mohamed Reda Bouadjenek
Missing data poses a significant challenge in data science, affecting decision-making processes and outcomes. Understanding what missing data is, how it occurs, and why it is crucial to handle it appropriately is paramount when working with real-world data, especially in tabular data, one of the most commonly used data…
Edoardo Costantini, Kyle M. Lang, Tim Reeskens, Klaas Sijtsma
Including a large number of predictors in the imputation model underlying a multiple imputation (MI) procedure is one of the most challenging tasks imputers face. A variety of high-dimensional MI techniques can help, but there has been limited research on their relative performance. In this study, we investigated a…
Panpan Zhang, Sharon X. Xie
In this paper, we compare the performance of available-case analysis (ACA) and several multiple imputation (MI) approaches for handling missing data problems in longitudinal analysis through estimation bias and relative efficiency. When the missingness of covariates depends on observed responses, ACA produces…
Paul T. von Hippel
Predictive mean matching (PMM) is a popular imputation strategy that imputes missing values by borrowing observed values from other cases with similar expectations. We show that, unlike other imputation strategies, PMM is not guaranteed to be consistent—and in fact can be severely biased—when values are missing at…
Yongseok Lee, Walter L. Leite
Researchers using propensity score analysis (PSA) to estimate treatment effects using secondary data may have to handle data that are missing not at random (MNAR). Existing methods for PSA with MNAR data use logistic regression to model the missing data mechanisms, thus requiring manual specification of functional…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Diba Behnoudfar, Cory Simon, Joshua Schrier
Aqueous, two-phase systems (ATPSs) may form upon mixing two solutions of independently water-soluble compounds. Many separation, purification, and extraction processes rely on ATPSs. Predicting the miscibility of solutions can accelerate and reduce the cost of the discovery of new ATPSs for these applications. Whereas…
Mithilesh Prakash, Jussi Tohka
We introduce a new subtype of ‘Missing Not at Random’ (MNAR) data, where the missingness is correlated with the labels (y) to be predicted, termed (y)-dependent MNAR. We demonstrate that this subtype can significantly bias the estimation of performance metrics in typical machine learning tasks. Unbiased error…
Sun Sun, Nan Luo, Erik Stenberg, Lars Lindholm + 6 more
'Karl A. Franklin' 'Yang Cao' 'Bian Liu' 'Lihua Li' 'Liangyuan Hu'] One of the main challenges for the successful implementation of health-related quality of life (HRQoL) assessments is missing data. The current study examined the feasibility and validity of a sequential multiple imputation (MI) method to deal with…
Barbora Rehák Bučková, Charlotte Fraza, Cecilie Koldbæk Lemvigh, Camilla Bärthel Flaaten + 11 more
Missing data remain a ubiquitous and critical challenge in large-scale clinical studies. Despite advances in imputation, most existing methods fail to address structured missingness, where data are missing according a deterministic pattern and which arise due to systematic patterns introduced by experimental design…
Authors not listed
Solute carrier (SLC) transporters constitute the largest family of membrane transport proteins in humans. They facilitate the movement of ions, neurotransmitters, nutrients, and drugs. Given their critical role in regulating cellular physiology, they are important therapeutic targets for neurological and psychological…
Tieu-Long Phan, Klaus Weinbauer, Thomas Gärtner, Daniel Merkle + 3 more
Purpose: Reaction databases are a key resource for a wide variety of applications in computational chemistry and biochemistry, including Computer-aided Synthesis Planning (CASP) and the large-scale analysis of metabolic networks. The full potential of these resources can only be realized if datasets are accurate and…
Mingze Bai, Jingwen Deng, Chengxin Dai, Julianus Pfeuffer + 1 more
Testing for significant differences in quantities on protein level is a common goal of many LFQ-based mass spectrometry proteomics experiments. Starting from a table of protein and/or peptide quantities from a fixed proteomics quantification software, there exists a multitude of tools and R packages to perform the…