26 papers · ranked by Valyu relevance
Amaze Lusompa
Model selection criteria are one of the most important tools in statistics. Proofs showing a model selection criterion is asymptotically optimal are tailored to the type of model (linear regression, quantile regression, penalized regression, etc.), the estimation method (linear smoothers, maximum likelihood…
Ryan Cecil, Lucas Mentch
Classical model selection seeks to find a single model within a particular class that optimizes some pre-specified criteria, such as maximizing a likelihood or minimizing a risk. More recently, there has been an increased interest in model set selection (MSS), where the aim is to identify a (confidence) set of…
Baidu Li, Xinhai Li
Linear models, including t-test, ANOVA, regression, ANCOVA, and generalized linear models, are foundational tools in statistical analysis. For large datasets, such as those involving tens of thousands of genes and millions of records, numerous advanced methods have been developed to improve both computational…
Alexandre René, André Longtin
Fitting models to data is an important part of the practice of science. Advances in machine learning have made it possible to fit more-and more complex-models, but have also exacerbated a problem: when multiple models fit the data equally well, which one(s) should we pick? The answer depends entirely on the modelling…
Prasanta S. Bandyopadhyay, Samidha Shetty, Gordon Brittan Jr., Geert Verdoolaege
An extensive literature on decision theory has been developed by both subjective Bayesians and Neyman-Pearson (NP) theorists, with more recent contributions to it from evidential decision theorists. The last-mentioned, however, have often been framed from a Bayesian perspective and therefore retain a subjectivist…
Michael C. Chung, Alen Zacharia, Juan Guan
Data-driven model discovery (DDMD) algorithms are powerful tools for extracting interpretable symbolic models from data. However, identifying the model that best balances goodness-of-fit and sparsity is often a laborious process requiring user fine-tuning, is prone to overfitting, and results may significantly vary…
L. M. André, J. L. Wadsworth, R. Huser
Likelihood-free approaches are appealing for performing inference on complex dependence models, either because it is not possible to formulate a likelihood function, or its evaluation is very computationally costly. This is the case for several models available in the multivariate extremes literature, particularly for…
By Riyadh Alrawkan, Edward Boone, Ryad Ghanam, Anton Westveld
Variable selection in linear regression models has been a problem since hypothesis testing began. Which variables to include or exclude from a model is not an easy task. Techniques such as Forward, Back ward, Stepwise Regression sequentially add or delete variables from a model. Penalized likelihood methods such as…
José Camacho
The validation of a data-driven model is the process of assessing the model's ability to generalize to new, unseen data in the population of interest. This paper proposes a set of general rules for model validation. These rules are designed to help practitioners create reliable validation plans and report their results…
Kenichiro McAlinn, Kōsaku Takanashi
Cross-validation is a standard technique used across science to test how well a model predicts new data. Data are split into K "folds," where one fold (i.e., hold-out set) is used to evaluate a model's predictive ability. Researchers typically rely on conventions when choosing K, commonly K = 5, or 80:20 split, even…
Zoran Levnajić
Understanding the processes behind the evolution of complex networks is a key objective in network science. An effective framework for tackling this challenge is network model selection, which involves finding the model from a set of candidates that best explains a given network. This book is a systematic review of…
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Charlotte Collingwood, Francesca Greenstreet, Marcus Stephenson-Jones, Rafal Bogacz
Action-selection is determined by a combination of goal-directed and habitual processes. Habits are defined as the reward-independent, stimulus-response relationships which form when an action is regularly executed in the same context, regardless of outcome. An influential computational model proposes that habit…
Francesco G. Rinaldi, Eugenio Piasini
To make sense of a noisy world, living beings constantly face decisions between competing interpretations for ambiguous sensory data. This process parallels statistical model selection, where most frameworks, like the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), are based on a…
Authors not listed
We present a multi-stage framework for predictive modeling that integrates automated feature engineering, selective dimensionality reduction, and targeted ensembling. Our pipeline begins with feature generation using a GPU-accelerated adaptation of AutoFeat, followed by variance-based pruning and LightGBM gain-based…
Grigoriy Gogoshin, Andrei S Rodin
Bayesian network (BN) modeling and computational systems biology have a long history of productive synergy. Learning BN structure from multiscale biomedical data is a central problem in this context. Computational methods for model reconstruction inherit the limitations of the underlying model selection criteria. As a…
Santiago Herce Castañon, Christopher R. Stephens
Predicting and understanding behaviour is a primary objective of many disciplines, especially human behaviour, as it is the cause of many of the world’s most pressing problems. Although it is a fundamental concept in multiple disciplines, there is no agreed operational definition of what it is. Neither is there a…
Kim, Dongseok, Choi, Hyoungsun + 4 more
We propose ϕ-test, a global feature-selection and significance procedure for black-box predictors that combines Shapley attributions with selective inference. Given a trained model and an evaluation dataset, ϕ-test performs SHAP-guided screening and fits a linear surrogate on the screened features via a selection rule…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Authors not listed
Machine learning potentials (MLPs) can help bridge the length- and time-scale gaps required to study diverse physicochemical phenomena in nanoporous materials with ab initio accuracy. These MLPs are typically trained on quantum chemical data obtained from traditional molecular dynamics (MD) simulations that…
Sina Kanannejad, Noemi Bongiorni, Elisa Nordera, Sara Redaelli + 4 more
Single-cell RNA sequencing has enabled the construction of comprehensive cell atlases, yet the quality and coherence of the cell-type annotations within these atlases remain largely unexamined. When a label is applied to a transcriptionally heterogeneous population, the downstream analyses that depend on it, and…
Authors not listed
DNA-encoded libraries (DELs) have emerged as a powerful platform for screening ultra-large chemical spaces by leveraging DNA barcodes to tag and track individual small molecules. Recent work has shown that machine learning can enhance DEL based hit discovery by denoising sequencing artifacts and improving binder…
Authors not listed
Incorporating prior domain knowledge into Bayesian optimization (BO) remains difficult for statistical methods, which also typically suffer from limited interpretability. Large language models (LLMs) offer complementary strengths in reasoning and knowledge integration, but it remains unclear when and how they improve…
Authors not listed
Active learning is an emerging paradigm used to help accelerating drug discovery, but most prior applications seek solely to optimize potency, whereas multiple properties influence a compound’s utility as a drug candidate. We introduce a method for multiobjective ligand optimization, which is able to efficiently handle…
Authors not listed
Machine learning (ML) models are increasingly used in quantum chemistry, but their reliability hinges on uncertainty quantification (UQ). In this study, we compare two prominent UQ paradigms—Deep Evidential Regression (DER) and Deep Ensembles—on the QM9 and WS22 datasets, with a specific emphasis on the role of post…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…