26 papers · ranked by Valyu relevance
Ben Klemens
This paper proposes a single form for statistical models that accommodates a broad range of models, from ordinary least squares to agent-based microsimulations. The definition makes it almost trivial to define morphisms to transform and combine existing models to produce new models. It offers a unified means of…
Andrew Gelman, Aki Vehtari, Daniel Simpson, Charles C. Margossian + 6 more
'Bob Carpenter' 'Yuling Yao' 'Lauren Kennedy' 'Jonah Gabry' 'Paul‐Christian Bürkner' 'Martin Modrák'] The Bayesian approach to data analysis provides a powerful way to handle uncertainty in all observations, model parameters, and model structure using probability theory. Probabilistic programming languages make it…
Joram Soch, Carsten Allefeld
We propose the statistical modelling approach to supervised learning (i.e. predicting labels from features) as an alternative to algorithmic machine learning (ML). The approach is demonstrated by employing a multivariate general linear model (MGLM) describing the effects of labels on features, possibly accounting for…
Juan Sosa, Lina Buitrago
We provide four case studies that use Bayesian machinery to making inductive reasoning. Our main motivation relies in offering several instances where the Bayesian approach to data analysis is exploited at its best to perform complex tasks, such as description, testing, estimation, and prediction. This work is not…
Andrew Gelman, Aki Vehtari
We review the most important statistical ideas of the past half century, which we categorize as: counterfactual causal inference, bootstrapping and simulation-based inference, overparameterized models and regularization, Bayesian multilevel models, generic computation algorithms, adaptive decision analysis, robust…
Zoubin Ghahramani
Modelling is fundamental to many fields of science and engineering. A model can be thought of as a representation of possible data one could predict from a system. The probabilistic approach to modelling uses probability theory to express all aspects of uncertainty in the model. The probabilistic approach is synonymous…
Michele Bennett, Karin Hayes, Ewa J. Kleczyk, Rajesh Mehta
Data scientists and statisticians are often at odds when determining the best approach – machine learning or statistical modeling – to solve an analytics challenge. However, machine learning and statistical modeling are more cousins than adversaries on different sides of an analysis battleground. Choosing between the…
Xuming He, David Madigan, Bin Yu, Jon Wellner
| EXECUTIVE SUMMARY | 4 | | --- | --- | | SECTION 1: ROLE/VALUE OF STATISTICS AND DATA SCIENCE | 6 | | SECTION 2: CHALLENGES IN SCIENTIFIC AND SOCIAL APPLICATIONS | 10 | | SECTION 3: FOUNDATIONAL RESEARCH | 16 | | SECTION 4: PROFESSIONAL CULTURE & COMMUNITY RESPONSIBILITIES | 20 | | SECTION 5: DOCTORAL EDUCATION | 23 |…
Joshua P. Jahner, C. Alex Buerkle, Dustin G. Gannon, Eliza M. Grames + 11 more
The proliferation of biological data with large numbers of samples and many dimensions is kindling hope that life scientists will be able to fit statistical and machine learning models that are highly predictive and interpretable. However, large biological data sets are commonly burdened with an inherent trade-off…
Danilo Bzdok, Denis Engemann, Olivier Grisel, Gaël Varoquaux + 1 more
In the 20^th^ century many advances in biological knowledge and evidence-based medicine were supported by p-values and accompanying methods. In the beginning 21^st^ century, ambitions towards precision medicine put a premium on detailed predictions for single individuals. The shift causes tension between traditional…
Teegwendé V. Porgo, Susan L. Norris, Georgia Salanti, Leigh F. Johnson + 4 more
'Leigh F. Johnson' 'Julie A. Simpson' 'Nicola Low' 'Matthias Egger' 'Christian L. Althaus'] Mathematical modeling studies are increasingly recognised as an important tool for evidence synthesis and to inform clinical and public health decision-making, particularly when data from systematic reviews of primary studies do…
Colin D. Kinz-Thompson, Korak Kumar Ray, Ruben L. Gonzalez
Biophysics experiments performed at single-molecule resolution contain exceptional insight into the structural details and dynamic behavior of biological systems. However, extracting this information from the corresponding experimental data unequivocally requires applying a biophysical model. Here, we discuss how to…
Christian Damgaard
In many applied cases of ecological and environmental modelling, there is a sizeable variation among the measured variables due to measurement- and sampling error. Such measurement- and sampling error among the independent variables may lead to regression dilution and biased prediction intervals in traditional…
Lihan Yan, Yongmin Sun, Michael R. Boivin, Paul O. Kwon + 1 more
This paper reviews several common challenges encountered in statistical analyses of epidemiological data for epidemiologists. We focus on the application of linear regression, multivariate logistic regression, and log-linear modeling to epidemiological data. Specific topics include: (a) deletion of outliers, (b)…
Andrew Gelman, Keith O’Rourke, Carlos Alberto De Bragança Pereira, Paulo Canas Rodrigues + 1 more
'Paulo Canas Rodrigues' 'Mark Andrew Gannon'] Amalgamation of evidence in statistics is conducted in several ways. Within a study, multiple observations are combined by averaging, or as factors in a likelihood or prediction algorithm. In multilevel modeling or Bayesian analysis, population or prior information is…
Robert Arbon, Yanchen Zhu, Antonia S. J. S. Mey
Markov state models (MSM) are a popular statistical method for analyzing the conformational dynamics of proteins, including protein folding. With all statistical and machine learning (ML) models choices must be made about the modeling pipeline that cannot be directly learned from the data. These choices, or…
Matti T. J. Heino, Matti Vuorre, Nelli Hankonen
Introduction Evaluating effects of behavior change interventions is a central interest in health psychology and behavioral medicine. Researchers in these fields routinely use frequentist statistical methods to evaluate the extent to which these interventions impact behavior and the hypothesized mediating processes in…
Authors not listed
Quantitative Structure-Activity Relationship (QSAR) modeling is a pillar of computational drug discovery. However, standard machine learning (ML) models are often confounded by the high-dimensional and intensely correlated nature of molecular descriptors. A model may identify a "bulk" property (e.g., molecular weight)…
Wenbiao Hu, Rebecca A. O'Leary, Kerrie Mengersen, Samantha Low Choy + 1 more
'Zheng Su'] Background Classification and regression tree (CART) models are tree-based exploratory data analysis methods which have been shown to be very useful in identifying and estimating complex hierarchical relationships in ecological and medical contexts. In this paper, a Bayesian CART model is described and…
Radu V. Craiu, Ruobin Gong, Xiao-Li Meng
This article proposes a set of categories, each one representing a particular distillation of important statistical ideas. Each category is labeled a "sense" because we think of these as essential in helping every statistical mind connect in constructive and insightful ways with statistical theory, methodologies, and…
James Wellnitz, Sankalp Jain, Joshua Hochuli, Travis Maxfield + 3 more
Traditional best practices for Quantitative Structure Activity Relationship (QSAR) modeling recommend dataset balancing and balanced accuracy (BA) as the key desired objective of model development. This study challenges the conventional norms by recommending the use of models with the highest positive predictive value…
Authors not listed
The rapid growth of worldwide computing power has transformed in silico chemistry into a discipline that is integrated into the daily work of many chemists. Nowadays, researchers find it increasingly straightforward to predict a wide range of molecular properties and chemi- cal processes at reasonable computational…
Authors not listed
Kinetic modeling is essential for predicting changes in food quality during processing and storage. This study evaluates the application of physics-informed neural networks (PINN) for food kinetic modeling, integrating kinetic insights into neural network frameworks. Based on three case studies, namely seed drying…
Roger W. Hoerl, Ronald D. Snee
Several authors, including the American Statistical Association (ASA), have noted the challenges facing statisticians when attacking large, complex and unstructured problems, as opposed to well-defined textbook problems. Clearly, the standard paradigm of selecting the one "correct" statistical method for such problems…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European Chemicals Agency (ECHA) and US Environmental Protection Agency (EPA) have listed approximately 800k chemicals that must be further investigated for their potential environmental and/or human health risk. A significant number of these chemicals have large enough global volumes of consumption (e.g.…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European and US chemical agencies have listed approximately 800k chemicals where knowledge on potential risks to human health and the environment are lacking. Filling these data gaps experimentally is impossible so in-silico approaches and prediction are essential. Many existing models are however limited by…