21 papers · ranked by Valyu relevance
Felipe Farias, Teresa B. Ludermir, Carmelo J. A. Bastos-Filho
The model selection procedure is usually a single-criterion decision making in which we select the model that maximizes a specific metric in a specific set, such as the Validation set performance. We claim this is very naive and can perform poor selections of over-fitted models due to the over-searching phenomenon…
Dilan Pathirana, Frank T. Bergmann, Domagoj Doresic, Polina Lakrisenko + 8 more
A central question in mathematical modeling of biological systems is determining which processes are relevant and how they can be described. There are often competing hypotheses, which yield different models. Model comparison requires parameter optimization and sampling methods. Yet, standards for the specification of…
Nicolas Lartillot
There is still no consensus as to how to select models in Bayesian phylogenetics, and more generally in applied Bayesian statistics. Bayes factors are often presented as the method of choice, yet other approaches have been proposed, such as cross-validation or information criteria. Each of these paradigms raises…
Dilan Pathirana, Frank T. Bergmann, Domagoj Doresic, Polina Lakrisenko + 8 more
A central question in mathematical modeling of biological systems is determining which processes are most relevant and how they can be described. There are often competing hypotheses, which yield different models. Model comparison requires parameter optimization and sampling methods. Yet, standards for the…
Baidu Li, Xinhai Li
Linear models, including t-test, ANOVA, regression, ANCOVA, and generalized linear models, are foundational tools in statistical analysis. For large datasets, such as those involving tens of thousands of genes and millions of records, numerous advanced methods have been developed to improve both computational…
Mohammad Ali Hajiani, Babak Seyfe
We propose a novel approach to select the best model of the data. Based on the exclusive properties of the nested models, we find the most parsimonious model containing the risk minimizer predictor. We prove the existence of probable approximately correct (PAC) bounds on the difference of the minimum empirical risk of…
Ryan Cecil, Lucas Mentch
Classical model selection seeks to find a single model within a particular class that optimizes some pre-specified criteria, such as maximizing a likelihood or minimizing a risk. More recently, there has been an increased interest in model set selection (MSS), where the aim is to identify a (confidence) set of…
Melissa Adrian, Jake A. Soloff, Rebecca Willett
Model selection is the process of choosing from a class of candidate models given data. For instance, methods such as the LASSO and sparse identification of nonlinear dynamics (SINDy) formulate model selection as finding a sparse solution to a linear system of equations determined by training data. However, absent…
Xinnong Li, Mark Sale, Keith Nieforth, James Craig + 6 more
Forward addition/backward elimination (FABE) has been the standard for population pharmacokinetic model selection (PPK) since NONMEM® was introduced. We investigated five machine learning (ML) algorithms (Genetic algorithm [GA], Gaussian process [GP], random forest [RF], gradient boosted random tree [GBRT], and…
Moayad Alnammi, Shengchao Liu, Spencer S Ericksen, Gene E Ananiev + 6 more
Traditional small molecule drug discovery is a time consuming and costly endeavor. High-throughput chemical screening can only assess a tiny fraction of drug-like chemical space. The strong predictive power of modern machine learning methods for virtual chemical screening enables training models on known active and…
Eugenio Piasini, Shuze Liu, Pratik Chaudhari, Vijay Balasubramanian + 1 more
Occam’s razor is the principle that, all else being equal, simpler explanations should be preferred over more complex ones^1^. This principle is thought to play a role in human perception and decision-making^2^, but the nature of our presumed preference for simplicity is not understood. Here we use preregistered…
Zihao Wen, David L. Dowe, Abhijit Mandal, Suneel Babu Chatla
Species distribution modeling is fundamental to biodiversity, evolution, conservation science, and the study of invasive species. Given environmental data and species distribution data, model selection techniques are frequently used to help identify relevant features. Existing studies aim to find the relevant features…
Hyemin Han
In the present study, I developed and tested an R module to explore the best models within the context of multilevel modeling in research in public health. The module that I developed, explore.models, compares all possible candidate models generated from a set of candidate predictors with information criteria, Akaike…
Yann McLatchie, Sölvi Rögnvaldsson, Frank Weber, Aki Vehtari
The concepts of Bayesian prediction, model comparison, and model selection have developed significantly over the last decade. As a result, the Bayesian community has witnessed a rapid growth in theoretical and applied contributions to building and selecting predictive models. Projection predictive inference in…
Kate E. Dray, Joseph J. Muldoon, Niall M. Mangan, Neda Bagheri + 1 more
Mathematical modeling is invaluable for advancing understanding and design of synthetic biological systems. However, the model development process is complicated and often unintuitive, requiring iteration on various computational tasks and comparisons with experimental data. Ad hoc model development can pose a barrier…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Xuan-Truc Dinh Tran, Tieu-Long Phan, Van-Thinh To, Ngoc-Vi Nguyen Tran + 4 more
3D pharmacophore models describe the ligand’s chemical interactions in their bioactive conformation. They offer a simple but sophisticated approach to decipher the chemically encoded ligand information, making them a valuable tool in Drug Design. Our research summarized the key studies for applying 3D pharmacophore…
Paul Francoeur, Daniel Penaherrera, David Koes
The immense size of chemical space, the relative scarcity of high quality data, and the cost of running experiments to accurately measure molecular properties makes active learning (AL) an attractive approach to efficiently explore the space and train high-quality models for molecular property prediction. While AL is…
Authors not listed
Machine learning holds significant promise for accelerating biomarker discovery in clinical proteomics, yet its real-world impact remains limited by widespread methodological pitfalls and unrealistic expectations. In this perspective, we critically examine the integration of machine learning into clinical proteomics…
Annette Spooner, Gelareh Mohammadi, Perminder S. Sachdev, Henry Brodaty + 1 more
'Henry Brodaty' 'Arcot Sowmya' ''] Background Feature selection is often used to identify the important features in a dataset but can produce unstable results when applied to high-dimensional data. The stability of feature selection can be improved with the use of feature selection ensembles, which aggregate the…