21 papers · ranked by Valyu relevance
Paul Madley-Dowd, Rachael A. Hughes, Maya B. Mathur, Jon Heron + 1 more
'Kate Tilling'] Missing data is a pervasive problem in epidemiology, with multiple imputation (MI) a commonly used analysis method. MI is valid when data are missing at random (MAR). However, definitions of MAR with multiple incomplete variables are not easily interpretable and graphical modelbased conditions are not…
Paul Madley-Dowd, Rachael A Hughes, Maya B Mathur, Jon Heron + 1 more
Missing data are a pervasive problem in epidemiology, with multiple imputation (MI) a commonly used analysis method. MI is valid when data are missing at random (MAR). However, definitions of MAR with multiple incomplete variables are not easily interpretable and descriptions of graphical model-based conditions are not…
A. Llera, M. Brammer, B. Oakley, J. Tillmann + 16 more
'J. S. Amelink' 'T. Mei' 'T. Charman' 'C. Ecker' 'F. Dell’Acqua' 'T. Banaschewski' 'C. Moessnang' 'S. Baron-Cohen' 'R. Holt' 'S. Durston' 'D. Murphy' 'E. Loth' 'J. K. Buitelaar' 'D. L. Floris' 'C. F. Beckmann'] An increasing number of large-scale multi-modal research initiatives has been conducted in the typically…
Xinru Wang, Lauren Kennedy, Qixuan Chen
The two-phase sampling design is a cost-effective sampling strategy that has been widely used in public health research. The conventional approach in this design is to create subsample specific weights that adjust for probability of selection and response in the second phase. However, these weights can be highly…
Dongyuan Song, Nan Miles Xi, Jingyi Jessica Li, Lin Wang
The number of cells measured in single-cell transcriptomic data has grown fast in recent years. For such large-scale data, subsampling is a powerful and often necessary tool for exploratory data analysis. However, the easiest random subsampling is not ideal from the perspective of preserving rare cell types. Therefore…
Jing Wang, Jiahui Zou, HaiYing Wang
Faced with massive data, subsampling is a commonly used technique to improve computational efficiency, and using nonuniform subsampling probabilities is an effective approach to improve estimation efficiency. For computational efficiency, subsampling is often implemented with replacement or through Poisson subsampling.…
Lin Wang
The use and analysis of massive data are challenging due to the high storage and computational cost. Subsampling algorithms are popular to downsize the data volume and reduce the computational burden. Existing subsampling approaches focus on data with numerical covariates. Although big data with categorical covariates…
Huiting Ou, Anuradha Surendra, Graeme S.V. McDowell, Emily Hashimoto-Roth + 3 more
Missing values are often unavoidable in modern high-throughput measurements due to various experimental or analytical reasons. Imputation, the process of replacing missing values in a dataset with estimated values, plays an important role in multivariate and machine learning analyses. Three missingness patterns have…
Amalan Mahendran, Helen Thompson, James M. McGree
Subsampling is a computationally efficient and scalable method to draw inference in large data settings based on a subset of the data rather than needing to consider the whole dataset. When employing subsampling techniques, a crucial consideration is how to select an informative subset based on the queries posed by the…
M. Templ, Markus Ulmer
Many imputation methods have been developed over the years and tested mostly under ideal settings. Surprisingly, there is no detailed research on how imputation methods perform when the idealized assumptions about the distribution of data and/or model assumptions are partly not fulfilled. This research looks into the…
Syed Abdul Rehman, Laila A. Al-Essa, Javid Shabbir, Zaheen Khan
Non-response is a common problem faced by surveyors while conducting surveys; this introduces a potential bias in the estimates of population parameters. One method of dealing with non-response is subsampling of the non-respondents, which increases precision in estimates by increasing the sample size. This study…
Vaishnavi Nagesh, Lauren Sanders, Sylvain V. Costes, Pinar Avci + 8 more
Missing data is a fundamental challenge in space biology, where high experimental costs, limited sample availability, and tissue allocation constraints produce datasets that are sparse, multimodal, and heterogeneous. We present a systematic four-stage framework for diagnosing, implementing, and validating data…
Tolou Shadbahr, Michael Roberts, Jan Stanczuk, Julian Gilbey + 14 more
'Philip Teare' 'Sören Dittmer' 'Matthew Thorpe' 'Ramon Viñas Torné' 'Evis Sala' 'Pietro Lió' 'Mishal Patel' 'Jacobus Preller' '' 'James H. F. Rudd' 'Tuomas Mirtti' 'Antti Sakari Rannikko' 'John A. D. Aston' 'Jing Tang' 'Carola-Bibiane Schönlieb'] Background Classifying samples in incomplete datasets is a common aim for…
Li Huang, Weikang Gong, Dongsheng Chen
Large single-cell RNA-sequencing (scRNA-seq) datasets offer unprecedented biological insights but pose major computational challenges for visualisation and analysis. Existing subsampling methods can improve efficiency yet may not guarantee downstream machine and deep learning (ML/DL) performance. Here, we propose…
Kimberly Conteddu, Prabhleen Kaur, Michael Brown, Julian Fennessy + 12 more
Animal populations are under mounting stress from the dual threats of climate change and rapid global human population growth, raising significant concerns about declining wildlife and the rising risk of zoonotic diseases. In many species, social interactions can be a highly plastic suite of behaviours that are…
Barbora Rehák Bučková, Charlotte Fraza, Cecilie Koldbæk Lemvigh, Camilla Bärthel Flaaten + 11 more
Missing data remain a ubiquitous and critical challenge in large-scale clinical studies. Despite advances in imputation, most existing methods fail to address structured missingness, where data are missing according a deterministic pattern and which arise due to systematic patterns introduced by experimental design…
Diba Behnoudfar, Cory Simon, Joshua Schrier
Aqueous, two-phase systems (ATPSs) may form upon mixing two solutions of independently water-soluble compounds. Many separation, purification, and extraction processes rely on ATPSs. Predicting the miscibility of solutions can accelerate and reduce the cost of the discovery of new ATPSs for these applications. Whereas…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Ayman Omar Baniamer, Henri Tilga
Statistical models are essential tools in data analysis. However, missing data plays a pivotal role in impacting the assumptions and effectiveness of statistical models, especially when there is a significant amount of missing data. This study addresses one of the core assumptions supporting many statistical models…
Authors not listed
Solute carrier (SLC) transporters constitute the largest family of membrane transport proteins in humans. They facilitate the movement of ions, neurotransmitters, nutrients, and drugs. Given their critical role in regulating cellular physiology, they are important therapeutic targets for neurological and psychological…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…