23 papers · ranked by Valyu relevance
Jean-Rémy Conti, Stéphan Clémençon
The ROC curve is the major tool for assessing not only the performance but also the fairness properties of a similarity scoring function. In order to draw reliable conclusions based on empirical ROC analysis, accurately evaluating the uncertainty level related to statistical versions of the ROC curves of interest is…
Christoph Dalitz, Felix Lögler
The m-out-of-n bootstrap is a possible workaround to compute confidence intervals for bootstrap inconsistent estimators, because it works under weaker conditions than the n-out-of-n bootstrap. It has the disadvantage, however, that it requires knowledge of an appropriate scaling factor τn and that the coverage…
Fumikazu Miwakeichi, Andreas Galka, Hector Zenil, Jiang Zhang + 1 more
'Peng Cui'] In this study, we present a thorough comparison of the performance of four different bootstrap methods for assessing the significance of causal analysis in time series data. For this purpose, multivariate simulated data are generated by a linear feedback system. The methods investigated are uncorrelated…
Jinji Pang, Wangqian Ju, Michael Welch, Phillip Gauger + 3 more
'Qijing Zhang' 'Chong Wang'] Developing and evaluating novel diagnostic assays are crucial components of contemporary diagnostic research. The receiver operating characteristic (ROC) curve and the area under the ROC curve (AUC) are frequently used to evaluate diagnostic assays’ performance. The variation in AUC…
Jared Clark, Richard L. Warr
Bootstrapping was designed to randomly resample data from a fixed sample using Monte Carlo techniques. However, the original sample itself defines a discrete distribution. Convolutional methods are well suited for discrete distributions, and we show the advantages of utilizing these techniques for bootstrapping. The…
Minghui Song, Guohua Zou, Alan T. K. Wan
Model averaging has gained significant attention in recent years due to its ability of fusing information from different models. The critical challenge in frequentist model averaging is the choice of weight vector. The bootstrap method, known for its favorable properties, presents a new solution. In this paper, we…
Muhammad Yahya Matdoan, Muhammad Mashuri, Muhammad Ahsan
Accurate parameter estimation is a critical component of effective process control using g charts. While traditional methods like maximum likelihood and Bayesian estimation are widely used, th ey may exhibit limitations in small sample size scenarios, leading to inaccurate parameter estimates. To address these…
Frédéric Lemoine, Olivier Gascuel
Felsenstein’s bootstrap is the most commonly used method to measure branch support in phylogenetics. Current sequencing technologies can result in massive sampling of taxa (e.g. SARS-CoV-2). In this case, the sequences are very close, the trees are short, and the branches correspond to a small number of mutations…
Amanda Forde, Gibran Hemani, John Ferguson
Genome-wide association studies (GWAS) are commonly used to identify genomic variants that are associated with complex traits, and estimate the magnitude of this association for each variant. However, it has been widely observed that the association estimates of variants tend to be lower in a replication study than in…
Guangming Li, Peida Zhan
Based on the standard of comparison and decision rules, the performance of these methods under different data conditions is graded and showed in [pone.0288069.t005], with the "+" symbol meaning accurate and the "-" symbol meaning inaccurate. As shown in [pone.0288069.t005], the traditional method, jackknife method and…
Yu Aikawa, Takeshi Morita, Kota Yoshimura
Recently, an application of the numerical bootstrap method to quantum mechanics was proposed, and it successfully reproduces the eigenstates of various systems. However, it is unclear why this method works. In order to understand this question, we study the bootstrap method in harmonic oscillators. We find that the…
Authors not listed
Plastic mechanical recycling is the conventional technological step towards circularity. In such aspects, complex mixtures of polyolefin blends are often fed into mechanical recycling systems, resulting in moulded products with uncertain quality. To add to the difficulty of heterogeneous feedstocks, the testing of…
Behnam Yousefimehr, Mehdi Ghatee, Mohammad Amin Seifi, Javad Fazli + 9 more
'Sajed Tavakoli' 'Zahra Rafei' 'Shervin Ghaffari' 'Abolfazl Nikahd' 'Mahdi Razi Gandomani' 'Alireza Orouji' 'Ramtin Mahmoudi Kashani' 'Sarina Heshmati' 'Negin Sadat Mousavi'] Imbalanced data poses a significant obstacle in machine learning, as an unequal distribution of class labels often results in skewed predictions…
Rory M. Crean, Joanna S. G. Slusky, Peter M. Kasson, Shina Caroline Lynn Kamerlin
Simulation datasets of proteins (e.g., those generated by molecular dynamics simulations) are filled with information about how the non-covalent interaction network within a protein regulates the conformation and thus function of said protein. Most proteins contain thousands of non-covalent interactions, with most of…
Wancen Mu, Eric Davis, Stuart Lee, Mikhail Dozmorov + 2 more
bootRanges provides fast functions for generation of bootstrapped genomic ranges representing the null sets in enrichment analysis. We show that shuffling or permutation schemes may result in overly narrow test statistics null distributions, while creating new ranges sets with a block bootstrap preserves local genomic…
Sandra Benítez-Peña, Rafael Blanquero, Emilio Carrizosa, Pepa Ramírez‐Cobo
'Pepa Ramírez‐Cobo'] Support vector machines (SVMs) are widely used and constitute one of the best examined and used machine learning models for two-class classification. Classification in SVM is based on a score procedure, yielding a deterministic classification rule, which can be transformed into a probabilistic rule…
Peter C. Austin
Background Healthcare provider profiling involves the comparison of outcomes between patients cared for by different healthcare providers. An important component of provider profiling is risk-adjustment so that providers that care for sicker patients are not unfairly penalized. One method for provider profiling entails…
Sukma Adi Perdana, Muhammad Mashuri, Muhammad Ahsan
This article proposes new method to improve the performance of bootstrap control chart for non-normal data. Bootstrap control charts for monitoring data require attention because the average run length (ARL) results of the bootstrap control charts can be less accurate and lacks stability. To deal with this issue, $X¯$…
Zeyuan Song, Sophia Gunn, Stefano Monti, Gina Marie Peloso + 3 more
Gaussian Graphical Models (GGM) have been widely used in biomedical research to explore complex relationships between many variables. There are well established procedures to build GGMs from a sample of independent and identical distributed observations. However, many studies include clustered and longitudinal data…
Sungmin Ji
In this study, gaussian mixture models with constrained parameter spaces are applied to the instar determination of insect species. Finite mixture models are often utilized to classify instars without knowing the instar number. Generally, parsimonious models with fewer free parameters would allow a more efficient…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Authors not listed
The rigorous design of adsorption-based separation processes, such as Pressure Swing Adsorption (PSA) and Temperature Swing Adsorption (TSA), is fundamentally dependent on the accuracy of the underlying mathematical models describing equilibrium isotherms and transport kinetics. However, the current state of the art is…
Robert Arbon, Yanchen Zhu, Antonia S. J. S. Mey
Markov state models (MSM) are a popular statistical method for analyzing the conformational dynamics of proteins, including protein folding. With all statistical and machine learning (ML) models choices must be made about the modeling pipeline that cannot be directly learned from the data. These choices, or…