22 papers · ranked by Valyu relevance
Fabian Klötzl, Bernhard Haubold, Alexander Bolshoy
We have recently developed a distance metric for efficiently estimating the number of substitutions per site between unaligned genome sequences. These substitution rates are called “anchor distances” and can be used for phylogeny reconstruction. Most phylogenies come with bootstrap support values, which are computed by…
Zhi Yang Tho, Raymond L. Chambers, A. H. Welsh
Clustered data arise naturally in many scientific and applied research settings where units are grouped within clusters. They are commonly analyzed using linear mixed models to account for within-cluster correlations. This article focuses on the scenario in which cluster sizes might be highly unbalanced and proposes a…
Elias Chaibub Neto, Kay Hamacher
In this paper we propose a vectorized implementation of the non-parametric bootstrap for statistics based on sample moments. Basically, we adopt the multinomial sampling formulation of the non-parametric bootstrap, and compute bootstrap replications of sample moment statistics by simply weighting the observed data…
F. Lemoine, J.-B. Domelevo Entfellner, E. Wilkinson, T. De Oliveira + 1 more
Felsenstein’s article describing the application of the bootstrap to evolutionary trees, is one of the most cited papers of all time. That statistical method, based on resampling and replications, is used extensively to assess the robustness of phylogenetic inferences. However, increasing numbers of sequences are now…
Frédéric Lemoine, Olivier Gascuel
Felsenstein’s bootstrap is the most commonly used method to measure branch support in phylogenetics. Current sequencing technologies can result in massive sampling of taxa (e.g. SARS-CoV-2). In this case, the sequences are very close, the trees are short, and the branches correspond to a small number of mutations…
Diep Thi Hoang, Olga Chernomor, Arndt von Haeseler, Bui Quang Minh + 1 more
The standard bootstrap (SBS), despite being computationally intensive, is widely used in maximum likelihood phylogenetic analyses. We recently proposed the ultrafast bootstrap approximation (UFBoot) to reduce computing time while achieving more unbiased branch supports than SBS under mild model violations. UFBoot has…
Julia Steinmetz, Carsten Jentsch
Mack's distribution-free chain ladder reserving model belongs to the most popular approaches in non-life insurance mathematics. Proposed to determine the first two moments of the reserve, it does not allow to identify the whole distribution of the reserve. For this purpose, Mack's model is usually equipped with a…
Joern H. Block, Christian Fisch, Mirko Hirschmann
Bootstrap financing refers to measures that entrepreneurial ventures undertake to preserve liquidity (e.g., reducing expenses, collecting receivables, delaying payments, preselling). Prior research shows that bootstrap financing is an important enabler for the growth of resource-constrained early-stage ventures.…
Joseph Rich, Lior Pachter
Summary: fastQpick is a command-line tool and Python library for sampling FASTQ reads with replacement. Sampling with replacement turns a single FASTQ file into an arbitrary number of bootstrap replicates, which enables uncertainty quantification and statistical analysis at the level of raw reads. This process answers…
Diep Thi Hoang, Le Sy Vinh, Tomáš Flouri, Alexandros Stamatakis + 2 more
'Arndt von Haeseler' 'Bui Quang Minh'] Background The nonparametric bootstrap is widely used to measure the branch support of phylogenetic trees. However, bootstrapping is computationally expensive and remains a bottleneck in phylogenetic analyses. Recently, an ultrafast bootstrap approximation (UFBoot) approach was…
Wancen Mu, Eric Davis, Stuart Lee, Mikhail Dozmorov + 2 more
bootRanges provides fast functions for generation of bootstrapped genomic ranges representing the null sets in enrichment analysis. We show that shuffling or permutation schemes may result in overly narrow test statistics null distributions, while creating new ranges sets with a block bootstrap preserves local genomic…
Rory M. Crean, Joanna S. G. Slusky, Peter M. Kasson, Shina Caroline Lynn Kamerlin
Simulation datasets of proteins (e.g., those generated by molecular dynamics simulations) are filled with information about how the non-covalent interaction network within a protein regulates the conformation and thus function of said protein. Most proteins contain thousands of non-covalent interactions, with most of…
Vladimir Makarenkov, Alix Boc, Jingxin Xie, Pedro Peres-Neto + 2 more
'François-Joseph Lapointe' 'Pierre Legendre'] Background Non-parametric bootstrapping is a widely-used statistical procedure for assessing confidence of model parameters based on the empirical distribution of the observed data [1] and, as such, it has become a common method for assessing tree confidence in…
Jia-Ming Chang, Evan W Floden, Javier Herrero, Olivier Gascuel + 3 more
ROC analyses are very useful to obtain an unbiased quantitative estimate of discriminative capacities, but they give little sense of a method usefulness in practical terms. In the case of a bootstrap support measure, the most important is to determine the key threshold value that makes it possible to distinguish…
Ian Osband, Benjamin Van Roy
This technical note presents a new approach to carrying out the kind of exploration achieved by Thompson sampling, but without explicitly maintaining or sampling from posterior distributions. The approach is based on a bootstrap technique that uses a combination of observed and artificially generated data. The latter…
Roman Matzutt, Jan Pennekamp, Erik Buchholz, Klaus Wehrle
Distributed anonymity services, such as onion routing networks or cryptocurrency tumblers, promise privacy protection without trusted third parties. While the security of these services is often well-researched, security implications of their required bootstrapping processes are usually neglected: Users either jointly…
Vincent Dufour-Decieux, Brandi Ransom, Rodrigo Freitas, Jose Blanchet + 1 more
Molecular Dynamics (MD) simulations are a key tool to understand the mechanism of complex chemical system and observe their outcomes in different conditions. However, such simulations are computationally expensive, which limits their timescales to the nanoseconds. This limitation is inconsequential at high…
Sihan Fu, Oucheng Liu, Shiyuan Wang, Jin Shi + 1 more
Code agents increasingly help developers work with unfamiliar repositories, but every such task depends on a costly prerequisite: bootstrapping the repository into a usable development state. This process requires substantial trial-and-error exploration, yet the resulting knowledge--resolved dependencies, repair…
Shriram Krishnamurthi, Emmanuel Schanzer, Joe Gibbs Politz, Benjamin S. Lerner + 2 more
'Benjamin S. Lerner' 'Kathi Fisler' 'Sam Dooman'] The Bootstrap Project's Data Science curriculum has trained about 100 teachers who are using it around the country. It is specifically designed to aid adoption at a wide range of institutions. It emphasizes valuable curricular goals by drawing on both the education…
Authors not listed
The rigorous design of adsorption-based separation processes, such as Pressure Swing Adsorption (PSA) and Temperature Swing Adsorption (TSA), is fundamentally dependent on the accuracy of the underlying mathematical models describing equilibrium isotherms and transport kinetics. However, the current state of the art is…
Priyanka Prakash Surve, Oleg Brodt, Mark Yampolskiy, Yuval Elovici + 1 more
'Asaf Shabtai'] Abstract—The Unified Extensible Firmware Interface (UEFI) is a linchpin of modern computing systems, governing secure system initialization and booting. This paper is urgently needed because of the surge in UEFI-related attacks and vulnerabilities in recent years. Motivated by this urgent concern, we…
Robert Arbon, Yanchen Zhu, Antonia S. J. S. Mey
Markov state models (MSM) are a popular statistical method for analyzing the conformational dynamics of proteins, including protein folding. With all statistical and machine learning (ML) models choices must be made about the modeling pipeline that cannot be directly learned from the data. These choices, or…