Search · four archives
Search · four archives
20 papers · ranked by Valyu relevance
David B. Dahl, Devin J. Johnson, R. Jacob Andros
Feature allocation models postulate a sampling distribution whose parameters are derived from shared features. Bayesian models place a prior distribution on the feature allocation, and Markov chain Monte Carlo is typically used for model fitting, which results in thousands of feature allocations sampled from the…
Alexandre Bouchard‐Côté, Andrew Roth
Bayesian feature allocation models are a popular tool for modelling data with a combinatorial latent structure. Exact inference in these models is generally intractable and so practitioners typically apply Markov Chain Monte Carlo (MCMC) methods for posterior inference. The most widely used MCMC strategies rely on an…
Lorenzo Ghilotti, Federico Camerlenghi, Tommaso Rigon
Feature allocation models are an extension of Bayesian nonparametric clustering models, where individuals can share multiple features. We study a broad class of models whose probability distribution has a product form, which includes the popular Indian buffet process. This class plays a prominent role among existing…
Mario Beraha, Federico Camerlenghi, Lorenzo Ghilotti
allocation models Authors: ['Mario Beraha' 'Federico Camerlenghi' 'Lorenzo Ghilotti'] We introduce and study a unified Bayesian framework for extended feature allocations which flexibly captures interactions – such as repulsion or attraction – among features and their associated weights. We provide a complete Bayesian…
Tamara Broderick, Jim Pitman, Michael I. Jordan
The problem of inferring a clustering of a data set has been the subject of much research in Bayesian analysis, and there currently exists a solid mathematical foundation for Bayesian approaches to clustering. In particular, the class of probability distributions over partitions of a data set has been characterized in…
Ali Foroughi pour, Lori A. Dalton
Background Many bioinformatics studies aim to identify markers, or features, that can be used to discriminate between distinct groups. In problems where strong individual markers are not available, or where interactions between gene products are of primary interest, it may be necessary to consider combinations of…
Aleksandra Vatian, Natalia Gusarova, Ivan Tomilov, Wei Li
In the modern world, there is a need to provide a better understanding of the importance or relevance of the available descriptive features for predicting target attributes to solve the feature ranking problem. Among the published works, the vast majority are devoted to the problems of feature selection and extraction…
Lai Jiang, Celia M. T. Greenwood, Weixin Yao, Longhai Li
Feature selection is demanded in many modern scientific research problems that use high-dimensional data. A typical example is to identify gene signatures that are related to a certain disease from high-dimensional gene expression data. The expression of genes may have grouping structures, for example, a group of…
Kaixin Yang, Long Liu, Yalu Wen
Feature selection is an indispensable step for the analysis of high-dimensional molecular data. Despite its importance, consensus is lacking on how to choose the most appropriate feature selection methods, especially when the performance of the feature selection methods itself depends on hyper-parameters. Bayesian…
Owen Forbes, Edgar Santos-Fernandez, Paul Pao-Yen Wu, Hong-Bo Xie + 7 more
'Paul E. Schwenn' 'Jim Lagopoulos' 'Lia Mills' 'Dashiell D. Sacks' 'Daniel F. Hermens' 'Kerrie Mengersen' 'Dariusz Siudak'] Various methods have been developed to combine inference across multiple sets of results for unsupervised clustering, within the ensemble clustering literature. The approach of reporting results…
Siying Li, Carol A. Seger, Meng Liu, Wenshan Dong + 2 more
In a dynamic environment, expectations of the future constantly change based on updated evidence and affect the dynamic allocation of attentional resources.To further investigate the neural mechanisms underlying efficient allocation of attention, we employed a modified Central Cue Posner paradigm in which the…
Aryan Deshwal, Cory Simon, Janardhan Rao Doppa
Given a gas storage or separation task, we wish to search a library of nanoporous materials (NPMs) for the one with the optimal adsorption property. The high cost of measuring the adsorption property of an NPM, whether in the lab or a simulation, precludes exhaustive search. We explain, demonstrate, and advocate…
Valentin Journé, Julien Papaïx, Emily Walker, François Courbet + 3 more
A trade-off between growth and fecundity, reflecting the inability of simultaneously investing in both functions when resources are limited, is a fundamental feature of life history theory. This particular trade-off is the result of evolutionary and environmental constrains shaping reproductive and growth traits, but…
Edward R Dougherty, Jianping Hua, Chao Sima
High-throughput biological technologies offer the promise of finding feature sets to serve as biomarkers for medical applications; however, the sheer number of potential features (genes, proteins, etc.) means that there needs to be massive feature selection, far greater than that envisioned in the classical literature.…
Haruo Hosoya, Aapo Hyvärinen
Although recent computational studies of feedforward neural network models have demonstrated remarkable performance in object recognition and neural response prediction, visual processing clearly has much more complex aspects that cannot be understood without feedback processing. Here, we propose a novel framework…
Sterling Baird, Jason R. Hall, Taylor D. Sparks
Would you rather search for a line inside a cube or a point inside a square? This type of solution degeneracy often exists in physics-based simulations and wet-lab experiments, but constraining these degeneracies is often unsupported or difficult to implement in many optimization packages, requiring additional time and…
Robert Arbon, Yanchen Zhu, Antonia S. J. S. Mey
Markov state models (MSM) are a popular statistical method for analyzing the conformational dynamics of proteins, including protein folding. With all statistical and machine learning (ML) models choices must be made about the modeling pipeline that cannot be directly learned from the data. These choices, or…
Michail Tsagris, Zacharias Papadovasilakis, Kleanthi Lakiotaki, Ioannis Tsamardinos
Feature selection seeks to identify a minimal-size subset of features that is maximally predictive of the outcome of interest. It is particularly important for biomarker discovery from high-dimensional molecular data, where the features could correspond to gene expressions, Single Nucleotide Polymorphisms (SNPs)…
Samir Rachid Zaim, Colleen Kenost, Joanne Berghout, Wesley Chiu + 3 more
In this era of data science-driven bioinformatics, machine learning research has focused on feature selection as users want more interpretation and post-hoc analyses for biomarker detection. However, when there are more features (i.e., transcript) than samples (i.e., mice or human samples) in a study, this poses major…
Shuji Shinohara, Nobuhito Manome, Kouta Suzuki, Ung-il Chung + 5 more
Bayesian inference is a process of narrowing down hypotheses (causes) to one that best explains observational data (effects). To accurately estimate a cause, a considerable amount of data is required to be observed for as long as possible. However, the object of inference is not always constant. In this case, a method…