21 papers · ranked by Valyu relevance
Zhi Yang Tho, Raymond L. Chambers, A. H. Welsh
Clustered data arise naturally in many scientific and applied research settings where units are grouped within clusters. They are commonly analyzed using linear mixed models to account for within-cluster correlations. This article focuses on the scenario in which cluster sizes might be highly unbalanced and proposes a…
Jiyue Qin, Samuel Davenport, Armin Schwartzman
Functional Magnetic Resonance Imaging (fMRI) is commonly used to localize brain regions activated during a task. Methods have been developed for constructing confidence regions of image excursion sets, allowing inference on brain regions exceeding non-zero activation thresholds. However, these methods have been limited…
Hanna Rajh-Weber, Stefan Ernest Huber, Martin Arendasy
Selecting an appropriate statistical method is a challenge frequently encountered by applied researchers, especially if assumptions for classical, parametric approaches are violated. To provide some guidelines and support, we compared classical hypothesis tests with their typical distributional assumptions of normality…
Esteban Charria-Girón, Laura Rosina Torres-Ortega, Joelle Mergola Greef, Yasmina Marin Felix + 4 more
Mass spectral molecular networking (MN) has emerged as a key computational approach to organize and analyze the vast volumes of tandem mass spectrometry (MS/MS) data generated in natural product research. MN connections are based on mass spectral similarities derived from cosine-based scores or machine learning-based…
Fredrik Lohne Aanes
In this paper I consider improving the KernelSHAP algorithm. I suggest to use the Wallenius' noncentral hypergeometric distribution for sampling the number of coalitions and perform sampling without replacement, so that the KernelSHAP estimation framework is improved further. I also introduce the Symmetric bootstrap to…
Minh Long Hoang, Cesare Svelto, Paolo Ciampolini, Guido Matrella + 3 more
This paper presents research on a Predictive and Uncertainty Assessment Framework (PUAF), providing a comparative analysis of two prominent methods, Monte Carlo (MC) Dropout and Bootstrap-based models, used in uncertainty estimation techniques of Neural Network predictions of human activity recognition using…
Chenghao Wei, Tianyu Zhang, Chen Li, Pukai Wang + 2 more
Tree-Augmented Naive Bayes (TAN) is an interpretable graphical structure model. However, its structure learning for continuous attributes depends on the class-conditional mutual information, which is sensitive to one-dimensional or two-dimensional density estimation. Accurate estimation is challenging under complex…
Lengyang Wang, Yingcun Xia, Ee Hui Goh, Mark Chen + 1 more
Timely detection of infectious disease outbreaks is critical for effective public health response. The effective reproduction number (R*t) is a key metric that captures transmission dynamics and signals the potential onset of outbreaks when it rises above 1. However, day-of-the-week and public holiday effects, along…
Ziwei Su, Imon Banerjee, Diego Klabjan
We propose and analyze a model-based bootstrap for transition kernels in finite controlled Markov chains (CMCs) with possibly nonstationary or history-dependent control policies, a setting that arises naturally in offline reinforcement learning (RL) when the behavior policy generating the data is unknown. We establish…
Sankalp Gilda
Finance, sensing, and demand streams violate the exchangeability that IID conformal prediction and the IID bootstrap assume, and existing libraries implement either a general resampling engine or conformal calibration without the other. tsbootstrap provides block, residual, sieve, and wild resampling, classical…
Sudhir Kumar, Koichiro Tamura, Sudip Sharma
Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting…
James O. McInerney, Christopher J. Creevey, Mary J. O’Connell
Phylogenetic inference relies on robust measures of branch support to assess the reliability of evolutionary relationships. While bootstrap resampling and Bayesian probabilities have become the predominant support metrics in molecular phylogenetics, decay indices (also known as Bremer support) provide an alternative…
Joseph Rich, Lior Pachter
Summary: fastQpick is a command-line tool and Python library for sampling FASTQ reads with replacement. Sampling with replacement turns a single FASTQ file into an arbitrary number of bootstrap replicates, which enables uncertainty quantification and statistical analysis at the level of raw reads. This process answers…
Xiaofeng Steven Liu
3## Bootstrap Bias Correction Bootstrap bias correction offers a simple solution to correcting the bias in the conditional ML estimate of odds ratio, that is used in Fisher's exact test. Bootstrap bias correction has two unique advantages over other alternative procedures. First, bootstrap does not make strong…
Gabriel Jimenez-Dominguez, Benjamin Audit, Pierre Borgnat, Patrice Ravel + 1 more
Understanding how gene regulatory networks respond to global cell perturbations remains a central challenge in systems biology and network inference. Modular Response Analysis (MRA) provides a mathematical framework to infer gene-to-gene directed connectivity graphs from perturbation experiments; however, classical MRA…
Jacques Raynal, Pierre Slangen, Elsa Raynal, Jacques Margerit
Observable performance is commonly used to characterize biological systems. In adaptive systems, however, similar performances may arise from distinct organizations, and configurations that appear comparable at a given time may follow different longitudinal trajectories. This limitation motivates a methodological…
Luke Honeybrook
Roughly half the cells in the human body are microbial, and changes in these communities are increasingly implicated in cardiovascular, metabolic, and oncological diseases. Yet identifying which taxa truly differ in abundance, differential abundance (DA), is distorted by four major sources of bias: loss of total…
Authors not listed
The rigorous design of adsorption-based separation processes, such as Pressure Swing Adsorption (PSA) and Temperature Swing Adsorption (TSA), is fundamentally dependent on the accuracy of the underlying mathematical models describing equilibrium isotherms and transport kinetics. However, the current state of the art is…
Authors not listed
Early prediction of drug-induced organ toxicity remains a major bottleneck in drug discovery and clinical pharmacotherapy. Most data-driven toxicity models behave as endpoint predictors: they output a label but provide limited transparency about why a compound is risky or which evidence channel dominated the decision.…
Authors not listed
Metal hydrides play a pivotal role in a wide range of applications, including hydrogen storage, compression, heat management, and catalysis, making them a central focus of interdisciplinary research spanning chemistry, materials science, and engineering. The performance of the metal hydride based systems is strongly…
Authors not listed
Quantitative Structure Activity Relationship (QSAR) remains an effective tool for early-stage chemical modelling and virtual screening in drug design. The advancements in this field are led by two core paradigms, 1) descriptor engineering, where complex fixed-length vectors of compounds are generated and conventional…