27 papers · ranked by Valyu relevance
Karl Friston, Will Penny
This note describes a Bayesian model selection or optimization procedure for post hoc inferences about reduced versions of a full model. The scheme provides the evidence (marginal likelihood) for any reduced model as a function of the posterior density over the parameters of the full model. It rests upon specifying…
Juho Piironen, Aki Vehtari
The goal of this paper is to compare several widely used Bayesian model selection methods in practical model selection problems, highlight their differences and give recommendations about the preferred approaches. We focus on the variable subset selection for regression and classification and perform several numerical…
Olha Bodnar, Viktor Eriksson
The location-scale model is usually present in physics and chemistry in connection to the Birge ratio method for the adjustment of fundamental physical constants such as the Planck constant or the Newtonian constant of gravitation, while the random effects model is the commonly used approach for meta-analysis in…
Colin D. Kinz-Thompson, Korak Kumar Ray, Ruben L. Gonzalez
Biophysics experiments performed at single-molecule resolution contain exceptional insight into the structural details and dynamic behavior of biological systems. However, extracting this information from the corresponding experimental data unequivocally requires applying a biophysical model. Here, we discuss how to…
Don van den Bergh, Merlise A. Clyde, Akash R. Komarlu Narendra Gupta, Tim de Jong + 4 more
'Tim de Jong' 'Quentin F. Gronau' 'Maarten Marsman' 'Alexander Ly' 'Eric-Jan Wagenmakers'] Linear regression analyses commonly involve two consecutive stages of statistical inquiry. In the first stage, a single ‘best’ model is defined by a specific selection of relevant predictors; in the second stage, the regression…
Jonathan H. Huggins, Jeffrey W. Miller
Bayesian model selection is premised on the assumption that the data are generated from one of the postulated models. However, in many applications, all of these models are incorrect (that is, there is misspecification). When the models are misspecified, two or more models can provide a nearly equally good fit to the…
A. P. Dawid, Monica Musio, Silvia Columbu
We consider the problem of choosing between parametric models for a discrete observable, taking a Bayesian approach in which the within-model prior distributions are allowed to be improper. In order to avoid the ambiguity in the marginal likelihood function in such a case, we apply a homogeneous scoring rule. For the…
Steven M Hill, Richard M Neve, Nora Bayani, Wen-Lin Kuo + 4 more
'Safiyyah Ziyad' 'Paul T Spellman' 'Joe W Gray' 'Sach Mukherjee'] Background An important question in the analysis of biochemical data is that of identifying subsets of molecular variables that may jointly influence a biological response. Statistical variable selection methods have been widely used for this purpose. In…
Michael Betancourt
As the frontiers of applied statistics progress through increasingly complex experiments we must exploit increasingly sophisticated inferential models to analyze the observations we make. In order to avoid misleading or outright erroneous inferences we then have to be increasingly diligent in scrutinizing the…
Henry Chai, Jean-François Ton, Roman Garnett, Michael A. Osborne
We present a novel technique for tailoring Bayesian quadrature (BQ) to model selection. The state-of-the-art for comparing the evidence of multiple models relies on Monte Carlo methods, which converge slowly and are unreliable for computationally expensive models. Previous research has shown that BQ offers sample…
Olivier Gimenez, Andy Royle, Marc Kéry, Chloé R. Nater + 1 more
In the hypothetico-deductive framework, models are compared to evaluate the relative strength of evidence in the data supporting alternative hypotheses. In Bayesian statistics, this is achieved by assessing models based on their probability of being true given the data, characterized by the posterior model probability.…
Nicolas Lartillot
There is still no consensus as to how to select models in Bayesian phylogenetics, and more generally in applied Bayesian statistics. Bayes factors are often presented as the method of choice, yet other approaches have been proposed, such as cross-validation or information criteria. Each of these paradigms raises…
Ang Li, Luis Pericchi, Kun Wang
There is not much literature on objective Bayesian analysis for binary classification problems, especially for intrinsic prior related methods. On the other hand, variational inference methods have been employed to solve classification problems using probit regression and logistic regression with normal priors. In this…
Richard Dybowski, Trevelyan J. McKinley, Pietro Mastroeni, Olivier Restif + 1 more
'Olivier Restif' 'Andrew J. Yates'] Understanding the mechanisms underlying the observed dynamics of complex biological systems requires the statistical assessment and comparison of multiple alternative models. Although this has traditionally been done using maximum likelihood-based methods such as Akaike's Information…
Yu Lin Hsu, Chu Chuan Jeng, Pavithra Sripathanallur Murali, Mohammadreza Torkjazi + 3 more
'Mohammadreza Torkjazi' 'J. S. West' 'Michaela Zuber' 'Vadim Sokolov'] This paper presents an overview of some of the concepts of Bayesian Learning. The number of scientific and industrial applications of Bayesian learning has been growing in size rapidly over the last few decades (Damien et al. 2013). This process has…
Bart van Erp, Wouter W. L. Nuijten, Thijs van de Laar, Bert de Vries + 1 more
'Patrick Shafto'] Bayesian state and parameter estimation are automated effectively in a variety of probabilistic programming languages. The process of model comparison on the other hand, which still requires error-prone and time-consuming manual derivations, is often overlooked despite its importance. This paper…
David J. Warne, Ruth E. Baker, Matthew J. Simpson
Reaction–diffusion models describing the movement, reproduction and death of individuals within a population are key mathematical modelling tools with widespread applications in mathematical biology. A diverse range of such continuum models have been applied in various biological contexts by choosing different flux and…
Eugenio Piasini, Shuze Liu, Pratik Chaudhari, Vijay Balasubramanian + 1 more
Occam’s razor is the principle that, all else being equal, simpler explanations should be preferred over more complex ones^1^. This principle is thought to play a role in human perception and decision-making^2^, but the nature of our presumed preference for simplicity is not understood. Here we use preregistered…
Robert Arbon, Yanchen Zhu, Antonia S. J. S. Mey
Markov state models (MSM) are a popular statistical method for analyzing the conformational dynamics of proteins, including protein folding. With all statistical and machine learning (ML) models choices must be made about the modeling pipeline that cannot be directly learned from the data. These choices, or…
Fábio K. Mendes, Remco Bouckaert, Luiz M. Carvalho, Alexei J. Drummond
Biology has become a highly mathematical discipline in which probabilistic models play a central role. As a result, research in the biological sciences is now dependent on computational tools capable of carrying out complex analyses. These tools must be validated before they can be used, but what is understood as…
Matti T. J. Heino, Matti Vuorre, Nelli Hankonen
Introduction Evaluating effects of behavior change interventions is a central interest in health psychology and behavioral medicine. Researchers in these fields routinely use frequentist statistical methods to evaluate the extent to which these interventions impact behavior and the hypothesized mediating processes in…
Authors not listed
Incorporating prior domain knowledge into Bayesian optimization (BO) remains difficult for statistical methods, which also typically suffer from limited interpretability. Large language models (LLMs) offer complementary strengths in reasoning and knowledge integration, but it remains unclear when and how they improve…
Aryan Deshwal, Cory Simon, Janardhan Rao Doppa
Given a gas storage or separation task, we wish to search a library of nanoporous materials (NPMs) for the one with the optimal adsorption property. The high cost of measuring the adsorption property of an NPM, whether in the lab or a simulation, precludes exhaustive search. We explain, demonstrate, and advocate…
Jonas Verhellen
In recent years, there have been considerable academic and industrial research efforts to develop novel generative models for high-performing, small molecules. Traditional, rules-based algorithms such as genetic algorithms [Jensen, Chem. Sci., 2019, 12, 3567-3572] have, however, been shown to rival deep learning…
Yannick Ureel, Maarten R. Dobbelaere, Yi Ouyang, Kevin De Ras + 3 more
By combining machine learning with design of experiments, so-called active machine learning, more efficient and cheaper research can be conducted. Machine learning algorithms are more flexible, and are better at investigating the processes spanning all length scales of chemical engineering. While the active machine…
Sterling Baird, Jason R. Hall, Taylor D. Sparks
Would you rather search for a line inside a cube or a point inside a square? Physics-based simulations and wet-lab experiments often have symmetries (degeneracies) that allow reducing problem dimensionality or search space, but constraining these degeneracies is often unsupported or difficult to implement in many…
Moayad Alnammi, Shengchao Liu, Spencer S Ericksen, Gene E Ananiev + 6 more
Traditional small molecule drug discovery is a time consuming and costly endeavor. High-throughput chemical screening can only assess a tiny fraction of drug-like chemical space. The strong predictive power of modern machine learning methods for virtual chemical screening enables training models on known active and…