Search · four archives
Search · four archives
23 papers · ranked by Valyu relevance
Soonil Kwon, Dai Wang, Xiuqing Guo
Genome-wide association studies usually involve several hundred thousand of single-nucleotide polymorphisms (SNPs). Conventional approaches face challenges when there are enormous number of SNPs but a relatively small number of samples and, in some cases, are not feasible. We introduce here an iterative Bayesian…
Gao Wang, Abhishek Sarkar, Peter Carbonetto, Matthew Stephens
We introduce a simple new approach to variable selection in linear regression, and to quantifying uncertainty in selected variables. The approach is based on a new model – the “Sum of Single Effects” (SuSiE) model – which comes from writing the sparse vector of regression coefficients as a sum of “single-effect”…
Nilotpal Sanyal
High-dimensional data with binary outcomes are common in biological sciences and related fields such as omics, epidemiology, and healthcare, environmental sciences, physical and engineering sciences, and business and policy making. Often, such data are sparse so that only a small number of all available variables truly…
Run Wang, Somak Dutta, Vivekananda Roy
Variable selection in ultra-high dimensional linear regression is often preceded by a screening step to significantly reduce the dimension. Here a Bayesian variable screening method (BITS) is developed. BITS can successfully integrate prior knowledge, if any, on effect sizes, and the number of true variables. BITS…
Wei Cheng, Sohini Ramachandran, Lorin Crawford
In this paper, we propose a new approach for variable selection using a collection of Bayesian neural networks with a focus on quantifying uncertainty over which variables are selected. Motivated by fine-mapping applications in statistical genetics, we refer to our framework as an “ensemble of single-effect neural…
Xitong Liang, Samuel Livingstone, Jim Griffin, Geert Verdoolaege + 1 more
'Ricardo Sandes Ehlers'] Developing an efficient computational scheme for high-dimensional Bayesian variable selection in generalised linear models and survival models has always been a challenging problem due to the absence of closed-form solutions to the marginal likelihood. The Reversible Jump Markov Chain Monte…
Dao Thanh Tung, Minh‐Ngoc Tran, Tran Manh Cuong
This article describes a full Bayesian treatment for simultaneous fixed-effect selection and parameter estimation in high-dimensional generalized linear mixed models. The approach consists of using a Bayesian adaptive Lasso penalty for signal-level adaptive shrinkage and a fast Variational Bayes scheme for estimating…
Kun Fan, Xiaoxi Li, Shejuty Devnath, Brock Olson + 2 more
Robust variable selection methods have emerged as powerful tools for dissecting high-dimensional gene-environment interactions in longitudinal studies, owing to their ability to accommodate intra-cluster correlations, capture structured sparsity, and handle heavy-tailed repeated measures. Despite these advantages…
Jacob Williams, Marco A. R. Ferreira, Tieming Ji
Background Single marker analysis (SMA) with linear mixed models for genome wide association studies has uncovered the contribution of genetic variants to many observed phenotypes. However, SMA has weak false discovery control. In addition, when a few variants have large effect sizes, SMA has low statistical power to…
Sierra A. Bainter, Thomas G. McCauley, Mahmoud M. Fahmy, Zachary T. Goodman + 2 more
'Zachary T. Goodman' 'Lauren B. Kupis' 'J. Sunil Rao'] In the current paper, we review existing tools for solving variable selection problems in psychology. Modern regularization methods such as lasso regression have recently been introduced in the field and are incorporated into popular methodologies, such as network…
Mahdi Nouraie, Connor Smith, Samuel Müller
Stability selection is a versatile framework for structure estimation and variable selection in high-dimensional setting, primarily grounded in frequentist principles. In this paper, we propose an enhanced methodology that integrates Bayesian analysis to refine the inference of inclusion probabilities within the…
Yong Li, Hefei Liu, Rubing Li, Lei Shi
Variable selection has always been an important issue in statistics. When a linear regression model is used to fit data, selecting appropriate explanatory variables that strongly impact the response variables has a significant effect on the model prediction accuracy and interpretation effect. redThis study introduces…
Federico Pavone, Juho Piironen, Paul‐Christian Bürkner, Aki Vehtari
Variable selection, or more generally, model reduction is an important aspect of the statistical workflow aiming to provide insights from data. In this paper, we discuss and demonstrate the benefits of using a reference model in variable selection. A reference model acts as a noise-filter on the target variable by…
Daniel F. Linder, Viral Panchal
In this paper we describe a Bayesian hierarchical model termed ‘PMMLogit’ for classification and model selection in high-dimensional settings with binary phenotypes as outcomes. Posterior computation in the logistic model is known to be computationally demanding due to its non-conjugacy with common priors. We combine a…
Yannick Ureel, Maarten R. Dobbelaere, Yi Ouyang, Kevin De Ras + 3 more
By combining machine learning with design of experiments, so-called active machine learning, more efficient and cheaper research can be conducted. Machine learning algorithms are more flexible, and are better at investigating the processes spanning all length scales of chemical engineering. While the active machine…
Anna Genell, Szilard Nemes, Gunnar Steineck, Paul W Dickman
Background Automatic variable selection methods are usually discouraged in medical research although we believe they might be valuable for studies where subject matter knowledge is limited. Bayesian model averaging may be useful for model selection but only limited attempts to compare it to stepwise regression have…
François Rousset, Raphaël Leblois, Arnaud Estoup, Jean-Michel Marin
Simulation-based methods such as approximate Bayesian computation (ABC) are widely used to infer the evolutionary history of populations from molecular genetic data. We describe and evaluate a new iterative method of statistical inference about model parameters, which revisits the idea of inferring a likelihood surface…
Junyang Qian, Wenfei Du, Yosuke Tanigawa, Matthew Aguirre + 3 more
Since its first proposal in statistics (1), the lasso has been an effective method for simultaneous variable selection and estimation. A number of packages have been developed to solve the lasso efficiently. However as large datasets become more prevalent, many algorithms are constrained by efficiency or memory bounds.…
Authors not listed
Modifying solution viscosity is a key functional application of polymers, yet the interplay of molecular chemistry, polymer architecture, and intermolecular interactions makes tailoring precise rheological responses challenging. We introduce a computational framework coupling topology-aware generative machine learning…
Authors not listed
Incorporating prior domain knowledge into Bayesian optimization (BO) remains difficult for statistical methods, which also typically suffer from limited interpretability. Large language models (LLMs) offer complementary strengths in reasoning and knowledge integration, but it remains unclear when and how they improve…
Ramsey Issa, Robert Sorenson, Taylor D. Sparks
The discovery of new dental materials is typically a slow process due to high-dimensionality of the formulation space as well as the multiple competing objectives which must be optimized for a given application. Here, we lay out a strategy using active learning and Bayesian optimization that has led to the discovery of…
Martin Robinson, Alan Bond, Alexandr Simonov, Jie Zhang + 1 more
Recently, we have introduced the use of techniques drawn from Bayesian statistics to recover kinetic and thermodynamic parameters from voltammetric data, and were able to show that the technique of large amplitude ac voltammetry yielded significantly more accurate parameter values than the equivalent dc approach. In…
Authors not listed
Solving optimization problems, especially for nonlinear and constrained systems, is a challenge. Decades of specialized algorithms have been developed for general and special cases of root finding, minimization (including constraints), for parameter estimation, and mapping connected spaces. These approaches typically…