23 papers · ranked by Valyu relevance
Yongshi Deng, Thomas Lumley
The use of multiple imputation (MI) is becoming increasingly popular for addressing missing data. Although some conventional MI approaches have been well studied and have shown empirical validity, they have limitations when processing large datasets with complex data structures. Their imputation performances usually…
Patrice L. Capers, Andrew W. Brown, John A. Dawson, David B. Allison
Background: Meta-research can involve manual retrieval and evaluation of research, which is resource intensive. Creation of high throughput methods (e.g., search heuristics, crowdsourcing) has improved feasibility of large meta-research questions, but possibly at the cost of accuracy. Objective: To evaluate the use of…
Paul Madley-Dowd, Rachael A. Hughes, Maya B. Mathur, Jon Heron + 1 more
'Kate Tilling'] Missing data is a pervasive problem in epidemiology, with multiple imputation (MI) a commonly used analysis method. MI is valid when data are missing at random (MAR). However, definitions of MAR with multiple incomplete variables are not easily interpretable and graphical modelbased conditions are not…
Paul Madley-Dowd, Rachael A Hughes, Maya B Mathur, Jon Heron + 1 more
Missing data are a pervasive problem in epidemiology, with multiple imputation (MI) a commonly used analysis method. MI is valid when data are missing at random (MAR). However, definitions of MAR with multiple incomplete variables are not easily interpretable and descriptions of graphical model-based conditions are not…
A. Llera, M. Brammer, B. Oakley, J. Tillmann + 16 more
'J. S. Amelink' 'T. Mei' 'T. Charman' 'C. Ecker' 'F. Dell’Acqua' 'T. Banaschewski' 'C. Moessnang' 'S. Baron-Cohen' 'R. Holt' 'S. Durston' 'D. Murphy' 'E. Loth' 'J. K. Buitelaar' 'D. L. Floris' 'C. F. Beckmann'] An increasing number of large-scale multi-modal research initiatives has been conducted in the typically…
Xinru Wang, Lauren Kennedy, Qixuan Chen
The two-phase sampling design is a cost-effective sampling strategy that has been widely used in public health research. The conventional approach in this design is to create subsample specific weights that adjust for probability of selection and response in the second phase. However, these weights can be highly…
Jianjun Hu, Haifeng Li, Michael S Waterman, Xianghong Jasmine Zhou
Background Missing value estimation is an important preprocessing step in microarray analysis. Although several methods have been developed to solve this problem, their performance is unsatisfactory for datasets with high rates of missing data, high measurement noise, or limited numbers of samples. In fact, more than…
Catrin O. Plumpton, Tim Morris, Dyfrig A. Hughes, Ian R. White
Background Missing data in a large scale survey presents major challenges. We focus on performing multiple imputation by chained equations when data contain multiple incomplete multi-item scales. Recent authors have proposed imputing such data at the level of the individual item, but this can lead to infeasibly large…
Huiting Ou, Anuradha Surendra, Graeme S.V. McDowell, Emily Hashimoto-Roth + 3 more
Missing values are often unavoidable in modern high-throughput measurements due to various experimental or analytical reasons. Imputation, the process of replacing missing values in a dataset with estimated values, plays an important role in multivariate and machine learning analyses. Three missingness patterns have…
Dongyuan Song, Nan Miles Xi, Jingyi Jessica Li, Lin Wang
The number of cells measured in single-cell transcriptomic data has grown fast in recent years. For such large-scale data, subsampling is a powerful and often necessary tool for exploratory data analysis. However, the easiest random subsampling is not ideal from the perspective of preserving rare cell types. Therefore…
Amalan Mahendran, Helen Thompson, James M. McGree
Subsampling is a computationally efficient and scalable method to draw inference in large data settings based on a subset of the data rather than needing to consider the whole dataset. When employing subsampling techniques, a crucial consideration is how to select an informative subset based on the queries posed by the…
M. Templ, Markus Ulmer
Many imputation methods have been developed over the years and tested mostly under ideal settings. Surprisingly, there is no detailed research on how imputation methods perform when the idealized assumptions about the distribution of data and/or model assumptions are partly not fulfilled. This research looks into the…
Jing Wang, Jiahui Zou, HaiYing Wang
Faced with massive data, subsampling is a commonly used technique to improve computational efficiency, and using nonuniform subsampling probabilities is an effective approach to improve estimation efficiency. For computational efficiency, subsampling is often implemented with replacement or through Poisson subsampling.…
Xikang Feng, Lingxi Chen, Zishuai Wang, Shuai Cheng Li
Single-cell RNA-sequencing (scRNA-seq) is essential for the study of cell-specific transcriptome landscapes. The scRNA-seq techniques capture merely a small fraction of the gene due to “dropout” events. When analyzing with scRNA-seq data, the dropout events receive intensive attentions. Imputation tools are proposed to…
Runmin Wei, Jingye Wang, Erik Jia, Tianlu Chen + 2 more
Left-censored missing values commonly exist in targeted metabolomics datasets and can be considered as missing not at random (MNAR). Improper data processing procedures for missing values will cause adverse impacts on subsequent statistical analyses. However, few imputation methods have been developed and applied to…
RA Hughes, JAC Sterne, K Tilling
Appropriate imputation inference requires both an unbiased imputation estimator and an unbiased variance estimator. The commonly used variance estimator, proposed by Rubin, can be biased when the imputation and analysis models are misspecified and/or incompatible. Robins and Wang proposed an alternative approach, which…
Kimberly Conteddu, Prabhleen Kaur, Michael Brown, Julian Fennessy + 12 more
Animal populations are under mounting stress from the dual threats of climate change and rapid global human population growth, raising significant concerns about declining wildlife and the rising risk of zoonotic diseases. In many species, social interactions can be a highly plastic suite of behaviours that are…
Barbora Rehák Bučková, Charlotte Fraza, Cecilie Koldbæk Lemvigh, Camilla Bärthel Flaaten + 11 more
Missing data remain a ubiquitous and critical challenge in large-scale clinical studies. Despite advances in imputation, most existing methods fail to address structured missingness, where data are missing according a deterministic pattern and which arise due to systematic patterns introduced by experimental design…
Diba Behnoudfar, Cory Simon, Joshua Schrier
Aqueous, two-phase systems (ATPSs) may form upon mixing two solutions of independently water-soluble compounds. Many separation, purification, and extraction processes rely on ATPSs. Predicting the miscibility of solutions can accelerate and reduce the cost of the discovery of new ATPSs for these applications. Whereas…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Authors not listed
Solute carrier (SLC) transporters constitute the largest family of membrane transport proteins in humans. They facilitate the movement of ions, neurotransmitters, nutrients, and drugs. Given their critical role in regulating cellular physiology, they are important therapeutic targets for neurological and psychological…
Seema Sangari, Herman E. Ray
Missing data is a common problem which has consistently plagued statisticians and applied analytical researchers. While replacement methods like mean-based or hot deck imputation have been well researched, emerging imputation techniques enabled through improved computational resources have had limited formal…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…