23 papers · ranked by Valyu relevance
Dongyue Xie, Wanrong Zhu, Matthew Stephens
We introduce a flexible empirical Bayes approach for fitting Bayesian generalized linear models. Specifically, we adopt a novel mean-field variational inference (VI) method and the prior is estimated within the VI algorithm, making the method tuning-free. Unlike traditional VI methods that optimize the posterior…
Zhanbolat Magzumov, Mustafa Kumral
The mining industry consumes about 1.7% of the energy generated worldwide, which is expected to increase in the coming decades. Milling is the most energy-intensive process of a typical mining operation. Many variables (e.g., rock characteristics, mineral matrix, and equipment properties) affect energy consumption.…
Satoko Hiura, Hiroki Abe, Kento Koyama, Shige Koseki
Conventional regression analysis using the least-squares method has been applied to describe bacterial behavior logarithmically. However, only the normal distribution is used as the error distribution in the least-squares method, and the variability and uncertainty related to bacterial behavior are not considered. In…
Zarina I. Vakhitova, Clair L. Alston-Knox, Yannick Griep
In the context of generalized linear models (GLMs), interactions are automatically induced on the natural scale of the data. The conventional approach to measuring effects in GLMs based on significance testing (e.g. the Wald test or using deviance to assess model fit) is not always appropriate. The objective of this…
Lucia Filippozzi, Iñigo Urteaga, Claudio Agostinelli
Covariate selection in Generalized Linear Models (GLMs) is a fundamental problem in statistics, as including irrelevant predictors might lead to overfitting and poor interpretability, while omitting relevant ones might result in biased estimates. Most Bayesian approaches to variable selection -- including…
Christos Argyropoulos, Andy P Grieve
Significance testing based on p-values has been implicated in the reproducibility crisis in scientific research, with one of the proposals being to eliminate them in favor of Bayesian analyses. Defenders of the p-values have countered that it is the improper use and errors in interpretation, rather than the pvalues…
Brian L. Trippe, Jonathan H. Huggins, Raj Agrawal, Tamara Broderick
Due to the ease of modern data collection, applied statisticians often have access to a large set of covariates that they wish to relate to some observed outcome. Generalized linear models (GLMs) offer a particularly interpretable framework for such an analysis. In these high-dimensional problems, the number of…
Ludger Starke, Dirk Ostwald
Variational Bayes (VB), variational maximum likelihood (VML), restricted maximum likelihood (ReML), and maximum likelihood (ML) are cornerstone parametric statistical estimation techniques in the analysis of functional neuroimaging data. However, the theoretical underpinnings of these model parameter estimation…
Ludger Starke, Dirk Ostwald
Variational Bayes (VB), variational maximum likelihood (VML), restricted maximum likelihood (ReML), and maximum likelihood (ML) are cornerstone parametric statistical estimation techniques in the analysis of functional neuroimaging data. However, the theoretical underpinnings of these model parameter estimation…
Joram Soch, Achim Meyer, John-Dylan Haynes, Carsten Allefeld
In functional magnetic resonance imaging (fMRI), model quality of general linear models (GLMs) for first-level analysis is rarely assessed. In recent work (32: “How to avoid mismodelling in GLM-based fMRI data analysis: cross-validated Bayesian model selection”, NeuroImage, vol. 141, pp. 469-489; DOI: 10.1016/j.…
Jack Sutton, Golnaz Shahtahmassebi, Quentin S. Hanley, Haroldo V. Ribeiro
'Haroldo V. Ribeiro'] Power law scaling models have been used to understand the complexity of systems as diverse as cities, neurological activity, and rainfall and lightning. In the scaling framework, power laws and standard linear regression methods are widely used to estimate model parameters with assumed normality…
Arthur Newbury
Estimating underlying cooccurrence relationships between pairs of species has long been a challenging task in ecology as the extent to which species actually cooccur is partially dependent on their prevalences. While recent work has taken large steps towards solving this problem, the next question is how to assess the…
Dao Thanh Tung, Minh‐Ngoc Tran, Tran Manh Cuong
This article describes a full Bayesian treatment for simultaneous fixed-effect selection and parameter estimation in high-dimensional generalized linear mixed models. The approach consists of using a Bayesian adaptive Lasso penalty for signal-level adaptive shrinkage and a fast Variational Bayes scheme for estimating…
Jie Pu, Di Fang, Jeffrey R. Wilson
Background The analysis of correlated binary data is commonly addressed through the use of conditional models with random effects included in the systematic component as opposed to generalized estimating equations (GEE) models that addressed the random component. Since the joint distribution of the observations is…
Samuel I. Berchuck, Felipe A. Medeiros, Sayan Mukherjee, Andréa Agazzi
'Andréa Agazzi'] The generalized linear mixed model (GLMM) is a popular statistical approach for handling correlated data, and is used extensively in applications areas where big data is common, including biomedical data settings. The focus of this paper is scalable statistical inference for the GLMM, where we define…
Joram Soch, Carsten Allefeld
In cognitive neuroscience, functional magnetic resonance imaging (fMRI) data are widely analyzed using general linear models (GLMs). However, model quality of GLMs for fMRI is rarely assessed, in part due to the lack of formal measures for statistical model inference. We introduce a new SPM toolbox for model…
Jeff S. Wesner, Justin P.F. Pomeranz
Bayesian data analysis is increasingly used in ecology, but prior specification remains focused on choosing non-informative priors (e.g., flat or vague priors). One barrier to choosing more informative priors is that priors must be specified on model parameters (e.g., intercepts, slopes, sigmas), but prior knowledge…
Lucian Chan, Geoffrey Hutchison, Garrett Morris
Generating low-energy molecular conformers is a key task for many areas of computational chemistry, molecular modeling and cheminformatics. Most current conformer generation methods primarily focus on generating geometrically diverse conformers rather than finding the most probable or energetically lowest minima. Here…
Lucian Chan, Geoffrey Hutchison, Garrett Morris
Generating low-energy molecular conformers is a key task for many areas of computational chemistry, molecular modeling and cheminformatics. Most current conformer generation methods primarily focus on generating geometrically diverse conformers rather than finding the most probable or energetically lowest minima. Here…
Yifan Wu, Aron Walsh, Alex Ganose
What is the minimum number of experiments, or calculations, required to find an optimal solution? Relevant chemical problems range from identifying a compound with target functionality within a given phase space to controlling materials synthesis and device fabrication conditions. A common feature in this application…
Jonas Verhellen
In recent years, there have been considerable academic and industrial research efforts to develop novel generative models for high-performing, small molecules. Traditional, rules-based algorithms such as genetic algorithms [Jensen, Chem. Sci., 2019, 12, 3567-3572] have, however, been shown to rival deep learning…
Authors not listed
Incorporating prior domain knowledge into Bayesian optimization (BO) remains difficult for statistical methods, which also typically suffer from limited interpretability. Large language models (LLMs) offer complementary strengths in reasoning and knowledge integration, but it remains unclear when and how they improve…
Sterling Baird, Jason R. Hall, Taylor D. Sparks
Would you rather search for a line inside a cube or a point inside a square? This type of solution degeneracy often exists in physics-based simulations and wet-lab experiments, but constraining these degeneracies is often unsupported or difficult to implement in many optimization packages, requiring additional time and…