23 papers · ranked by Valyu relevance
Yiran Dong, Chao-Ying Joanne Peng
The impact of missing data on quantitative research can be serious, leading to biased estimates of parameters, loss of information, decreased statistical power, increased standard errors, and weakened generalizability of findings. In this paper, we discussed and demonstrated three principled missing data methods…
Tlamelo Emmanuel, Thabiso Maupong, Dimane Mpoeleng, Thabo Semong + 2 more
'Banyatsang Mphago' 'Oteng Tabona'] Machine learning has been the corner stone in analysing and extracting information from data and often a problem of missing values is encountered. Missing values occur because of various factors like missing completely at random, missing at random or missing not at random. All these…
Gunther Eysenbach, Filip Smit, David Streiner, Matthijs Blankers + 2 more
'Maarten W J Koeter' 'Gerard M Schippers'] Background Missing data is a common nuisance in eHealth research: it is hard to prevent and may invalidate research findings. Objective In this paper several statistical approaches to data “missingness” are discussed and tested in a simulation study. Basic approaches (complete…
Pegah Golchian, Jan Kapar, David S. Watson, Marvin N. Wright
Handling missing values is a common challenge in biostatistical analyses, typically addressed by imputation methods. We propose a novel, fast, and easy-to-use imputation method called missing value imputation with adversarial random forests (MissARF), based on generative machine learning, that provides both single and…
Huiting Ou, Anuradha Surendra, Graeme S.V. McDowell, Emily Hashimoto-Roth + 3 more
Missing values are often unavoidable in modern high-throughput measurements due to various experimental or analytical reasons. Imputation, the process of replacing missing values in a dataset with estimated values, plays an important role in multivariate and machine learning analyses. Three missingness patterns have…
Huiting Ou, Anuradha Surendra, Graeme S V McDowell, Emily Hashimoto-Roth + 4 more
Missing data are a major problem for multivariate, machine learning (ML) and network analyses. For example, in large lipidomic or metabolic datasets, measurements for some analytes may not be available in every sample due to routine technical variability, low abundance, ion suppression from co-eluting analytes…
Robert Thiesmeier, Matteo Bottai, Nicola Orsini
Missing data is a common challenge across scientific disciplines. Current imputation methods require the availability of individual data to impute missing values. Often, however, missingness requires using external data for the imputation. In this paper, we introduce a new Stata command, mi impute from, designed to…
Ayman Omar Baniamer, Henri Tilga
Statistical models are essential tools in data analysis. However, missing data plays a pivotal role in impacting the assumptions and effectiveness of statistical models, especially when there is a significant amount of missing data. This study addresses one of the core assumptions supporting many statistical models…
Wei Qiu, Yangsibo Huang, Quanzheng Li
—Missing value imputation is a challenging and wellresearched topic in data mining. In this paper, we propose IFGAN, a missing value imputation algorithm based on Featurespecific Generative Adversarial Networks (GAN). Our idea is intuitive yet effective: a feature-specific generator is trained to impute missing values…
Youran Zhou, Sunil Aryal, Mohamed Reda Bouadjenek
Missing data poses a significant challenge in data science, affecting decision-making processes and outcomes. Understanding what missing data is, how it occurs, and why it is crucial to handle it appropriately is paramount when working with real-world data, especially in tabular data, one of the most commonly used data…
Seema Sangari, Herman E. Ray
Missing data is a common problem which has consistently plagued statisticians and applied analytical researchers. While replacement methods like mean-based or hot deck imputation have been well researched, emerging imputation techniques enabled through improved computational resources have had limited formal…
Luis Alejandro Masmela-Caita, Thais P. Galletti, Marcos O. Prates
Missing data theory deals with the statistical methods in the occurrence of missing data. Missing data occurs when some values are not stored or observed for variables of interest. However, most of the statistical theory assumes that data is fully observed. An alternative to deal with incomplete databases is to fill in…
Brett K. Beaulieu-Jones, Daniel R. Lavage, John W. Snyder, Jason H. Moore + 2 more
Missing data is a challenge for all studies; however, this is especially true for electronic health record (EHR) based analyses. Failure to appropriately consider missing data can lead to biased results. Here, we provide detailed procedures for when and how to conduct imputation of EHR data. We demonstrate how the…
Maria Thurow, Florian Dumpert, Burim Ramosaj, Markus Pauly
In statistical survey analysis, (partial) non-responders are integral elements during data acquisition. Treating missing values during data preparation and data analysis is therefore a non-trivial underpinning. Focusing on different data sets from the Federal Statistical Office of Germany (DESTATIS), we investigate…
Yongseok Lee, Walter L. Leite
Researchers using propensity score analysis (PSA) to estimate treatment effects using secondary data may have to handle data that are missing not at random (MNAR). Existing methods for PSA with MNAR data use logistic regression to model the missing data mechanisms, thus requiring manual specification of functional…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Diba Behnoudfar, Cory Simon, Joshua Schrier
Aqueous, two-phase systems (ATPSs) may form upon mixing two solutions of independently water-soluble compounds. Many separation, purification, and extraction processes rely on ATPSs. Predicting the miscibility of solutions can accelerate and reduce the cost of the discovery of new ATPSs for these applications. Whereas…
Mithilesh Prakash, Jussi Tohka
We introduce a new subtype of ‘Missing Not at Random’ (MNAR) data, where the missingness is correlated with the labels (y) to be predicted, termed (y)-dependent MNAR. We demonstrate that this subtype can significantly bias the estimation of performance metrics in typical machine learning tasks. Unbiased error…
Barbora Rehák Bučková, Charlotte Fraza, Cecilie Koldbæk Lemvigh, Camilla Bärthel Flaaten + 11 more
Missing data remain a ubiquitous and critical challenge in large-scale clinical studies. Despite advances in imputation, most existing methods fail to address structured missingness, where data are missing according a deterministic pattern and which arise due to systematic patterns introduced by experimental design…
Authors not listed
Solute carrier (SLC) transporters constitute the largest family of membrane transport proteins in humans. They facilitate the movement of ions, neurotransmitters, nutrients, and drugs. Given their critical role in regulating cellular physiology, they are important therapeutic targets for neurological and psychological…
Charles F. Manski
Incomplete observability of data generates an identification problem. There is no panacea for missing data. What one can learn about a population parameter depends on the assumptions one finds credible to maintain. The credibility of assumptions varies with the empirical setting. No specific assumptions can provide a…
Tieu-Long Phan, Klaus Weinbauer, Thomas Gärtner, Daniel Merkle + 3 more
Purpose: Reaction databases are a key resource for a wide variety of applications in computational chemistry and biochemistry, including Computer-aided Synthesis Planning (CASP) and the large-scale analysis of metabolic networks. The full potential of these resources can only be realized if datasets are accurate and…
Mingze Bai, Jingwen Deng, Chengxin Dai, Julianus Pfeuffer + 1 more
Testing for significant differences in quantities on protein level is a common goal of many LFQ-based mass spectrometry proteomics experiments. Starting from a table of protein and/or peptide quantities from a fixed proteomics quantification software, there exists a multitude of tools and R packages to perform the…