21 papers · ranked by Valyu relevance
Cheng Peng
Bootstrap resampling is the foundation of many ensemble learning methods, and Out-of-Bag (OOB) error estimation is by far the most popular way to get an internal assessment of generalization performance. In the standard multinomial bootstrap, the number of distinct observations in each resample is random. Although this…
Di Zhang
Bootstrapping is a powerful statistical resampling technique for estimating the sampling distribution of an estimator. However, its computational cost becomes prohibitive for large datasets or a high number of resamples. This paper presents a theoretical analysis and design of parallel bootstrapping algorithms using…
Sudhir Kumar, Koichiro Tamura, Sudip Sharma
Long runtime, high memory demands, and reliance on high-performance computing increasingly limit the evolutionary analysis of long phylogenomic datasets. We review a scalable framework based on phylogenomic subsampling and upsampling (PSU), in which many small subsamples of sites from a long concatenated sequence…
Johannes Bleher, Claudia Tarantola
When variable selection methods are applied to bootstrapped and multiply imputed datasets, the set of selected variables typically varies across iterations. Aggregating results via the union rule can lead to overly dense models. We propose a sequential evidence aggregation procedure that models detection outcomes…
Sudhir Kumar, Koichiro Tamura, Sudip Sharma
Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting…
Hyunwook Koh
Random Forest is a widely used tree-based ensemble learning algorithm that efficiently captures complex nonlinear relationships and higher-order feature interactions with no distributional assumptions to be satisfied. It is also well-suited to human microbiome studies, where the data are highly skewed, overdispersed…
Hanna Rajh-Weber, Stefan Ernest Huber, Martin Arendasy
Selecting an appropriate statistical method is a challenge frequently encountered by applied researchers, especially if assumptions for classical, parametric approaches are violated. To provide some guidelines and support, we compared classical hypothesis tests with their typical distributional assumptions of normality…
Ana Helena Tavares, Ana Silva, Tiago Freitas, Maria Costa + 3 more
Despite the advances on data analysis methodologies in the last decades, most of the traditional regression methods cannot be directly applied to large-scale data. Although aggregation methods are especially designed to deal with large-scale data, their performance may be strongly reduced in ill-conditioned problems…
Joseph Rich, Lior Pachter
Summary: fastQpick is a command-line tool and Python library for sampling FASTQ reads with replacement. Sampling with replacement turns a single FASTQ file into an arbitrary number of bootstrap replicates, which enables uncertainty quantification and statistical analysis at the level of raw reads. This process answers…
Timothy Christensen, Silvia Goncalves, Benoit Perron
AI/ML methods are increasingly used in economics to generate binary variables (or labels) via classification algorithms. When these generated variables are included as covariates in regressions, even small misclassification errors can induce large biases in OLS estimators and invalidate standard inference. We study…
Li Shandross, Emily Howerton, Lucie Contamin, Harry Hochheiser + 20 more
Combining predictions from multiple models into an ensemble is a widely used practice across many fields with demonstrated performance benefits. Popularized through domains such as weather forecasting and climate modeling, multi-model ensembles are becoming increasingly common in public health and biological…
Authors not listed
The rigorous design of adsorption-based separation processes, such as Pressure Swing Adsorption (PSA) and Temperature Swing Adsorption (TSA), is fundamentally dependent on the accuracy of the underlying mathematical models describing equilibrium isotherms and transport kinetics. However, the current state of the art is…
Authors not listed
Next Generation Risk Assessment (NGRA) promotes animal-free, exposure-informed, and hypothesis-driven approaches to chemical safety assessment. In silico tools, such as quantitative structure-activity relationship (QSAR) models, are valuable new approach methodologies (NAMs) for use in NGRA. However, the practical…
Fabian Stricker, Jose A. Peregrina, David Bermbach, Christian Zirpins
Performance evaluation is essential for assessing the quality of machine learning (ML) models and guiding deployment decisions. In federated learning (FL), assessing the performance is challenging because data are distributed across participants. Consequently, the coordinator must rely on locally computed evaluation…
Stephen McCoy, Daniel McBride, D. Katie McCullough, Benjamin C. Calfee + 3 more
We develop and apply a learning framework for parameter estimation in initial value problems that are assessed only indirectly via aggregate data such as sample means and/or standard deviations. Our comprehensive framework follows Bayesian principles and consists of specialized Markov chain Monte Carlo computational…
Rubén Grillo-Risco, Maksym Kupchyk Tiurin, Carla Perpiñá-Clérigues, Francisco J. Cordero Felipe + 3 more
The growing number of omics datasets in public repositories provides an opportunity to enhance data reusability through data integration; however, complex statistical barriers often hinder the effective combination of independent studies. To address this problem, we present MetaOmixTools, an interactive web-based suite…
Tomi Suomi, Jalmari Kettunen, Taneli Pusa, Laura L. Elo
Reproducibility is fundamental to reliable scientific discoveries. The reproducibility-optimized test statistic (ROTS) is a robust framework designed to identify reproducible features (e.g. genes or proteins) in high-dimensional differential expression analyses such as transcriptomics and proteomics. This is achieved…
Authors not listed
Early prediction of drug-induced organ toxicity remains a major bottleneck in drug discovery and clinical pharmacotherapy. Most data-driven toxicity models behave as endpoint predictors: they output a label but provide limited transparency about why a compound is risky or which evidence channel dominated the decision.…
Fook C. Cheong, Seung Y. Lee, Sujata Bais, Satyam Khanal + 1 more
Biomolecular condensates formed through phase separation are fundamental to cellular organization. Although the physical principles underlying intrinsically disordered proteins are well understood, the molecular determinants of condensate formation in globular proteins remain elusive. Here, we employ Holographic…
Authors not listed
Phase equilibrium calculations are crucial in chemical engineering design and optimization processes. The PC-SAFT equation of state (EoS) can precisely calculate phase equilibrium, but is relatively complex and computationally intensive. Surrogate models are mathematically simple models that map or regress the…
Heranga K. Rathnasekara, Sinjini Sikdar
Analysis of genomics data for predicting disease outcomes is a fast-growing field in medical research. There often exist categorical, specifically, ordinal outcomes that need to be predicted based on genomic profiles. This has led to recent development of some high-dimensional ordinal classification methods that can…