Search · four archives
Search · four archives
27 papers · ranked by Valyu relevance
Hu Shi, Nicola Pezzotti, Max Welling
In healthcare applications, predictive uncertainty has been used to assess predictive accuracy. In this paper, we demonstrate that predictive uncertainty estimated by the current methods does not highly correlate with prediction error by decomposing the latter into random and systematic errors, and showing that the…
David J. Wood, Lars Carlsson, Martin Eklund, Ulf Norinder + 1 more
'Jonna Stålring'] We propose that quantitative structure-activity relationship (QSAR) predictions should be explicitly represented as predictive (probability) distributions. If both predictions and experimental measurements are treated as probability distributions, the quality of a set of predictive distributions…
Jin Li, Qin Zhang
Assessing the accuracy of predictive models is critical because predictive models have been increasingly used across various disciplines and predictive accuracy determines the quality of resultant predictions. Pearson product-moment correlation coefficient (r) and the coefficient of determination (r2) are among the…
James E. Smith, Praminda Caleb-Solly, Muhammad Atif Tahir, Davy Sannen + 1 more
'Davy Sannen' 'Hendrik Van Brussel'] The accuracy of machine learning systems is a widely studied research topic. Established techniques such as cross-validation predict the accuracy on unseen data of the classifier produced by applying a given learning method to a given training data set. However, they do not predict…
David Anderson, Margret Bjarnadottir, Nebojsa Bacanin
How much information does a dataset contain about an outcome of interest? To answer this question, estimates are generated for a given dataset, representing the minimum possible absolute prediction error for an outcome variable that any model could achieve. The estimate is produced using a constrained omniscient model…
J. Emmanuel Johnson, Valero Laparra, Gustau Camps‐Valls
Gaussian Processes (GPs) are a class of kernel methods that have shown to be very useful in geoscience applications. They are widely used because they are simple, flexible and provide very accurate estimates for nonlinear problems, especially in parameter retrieval. An addition to a predictive mean function, GPs come…
Yi-Fang Hsu, Florian Waszak, Jarmo A. Hämäläinen
The predictive coding model of perception proposes that successful representation of the perceptual world depends upon cancelling out the discrepancy between prediction and sensory input (i.e., prediction error). Recent studies further suggest a distinction between prediction error associated with non-predicted stimuli…
Kei Irie, Phillip Minar, Jack Reifenberg, Brendan M Boyle + 3 more
Population pharmacokinetic (PK) model-based Bayesian estimation is widely used for dose individualization, particularly when sample availability is limited. However, its predictive accuracy can be compromised by factors such as misspecified prior information, intra-patient variability, and uncertainties in PK…
Christopher Meek
Understanding prediction errors and determining how to fix them is critical to building effective predictive systems. In this paper, we delineate four types of prediction errors (mislabeling, representation, learner and boundary errors) and demonstrate that these four types characterize all prediction errors. In…
Megan T. Jones, Ishaan Gadiyar, Xinyu Zhang, Kaidi Kang + 6 more
Machine learning is used in neuroscience to examine brain-phenotype associations and facilitate individual prediction from high-dimensional brain imaging. For continuous phenotypes, Pearson’s correlation between the observed and predicted phenotype is used to quantify model accuracy in testing data. However, recent…
Stefan Siegert, Philip G. Sansom, Robin M. Williams
Ensemble forecasts of weather and climate are subject to systematic biases in the ensemble mean and variance, leading to inaccurate estimates of the forecast mean and variance. To address these biases, ensemble forecasts are post-processed using statistical recalibration frameworks. These frameworks often specify…
Laura C Rosella, Paul Corey, Therese A Stukel, Cam Mustard + 2 more
'Doug G Manuel'] Background Self-reported height and weight are commonly collected at the population level; however, they can be subject to measurement error. The impact of this error on predicted risk, discrimination, and calibration of a model that uses body mass index (BMI) to predict risk of diabetes incidence is…
Siruo Wang, Tyler H. McCormick, Jeffrey T. Leek
Many modern problems in medicine and public health leverage machine learning methods to predict outcomes based on observable covariates [1, 2, 3, 4]. In an increasingly wide array of settings, these predicted outcomes are used in subsequent statistical analysis, often without accounting for the distinction between…
Jiying Yan, Charles Rahal
For Correspondence: Jiani Yan and Charles Rahal, Leverhulme Centre for Demographic Science, Demographic Science Unit, University of Oxford, OX1 1JD, United Kingdom. Tel: 01865 286170. Email: [jiani.yan@sociology.ox.ac.uk](mailto:jiani.yan@sociology.ox.ac.uk) and…
Authors not listed
Accurate determination of the metabolic fate of xenobiotics is essential for ensuring their safety and efficacy. While in vivo and in vitro methods remain the gold standard for assessing metabolic properties, they are both costly and time-consuming. In silico metabolism prediction models offer complementary solutions…
Mithilesh Prakash, Jussi Tohka
We introduce a new subtype of ‘Missing Not at Random’ (MNAR) data, where the missingness is correlated with the labels (y) to be predicted, termed (y)-dependent MNAR. We demonstrate that this subtype can significantly bias the estimation of performance metrics in typical machine learning tasks. Unbiased error…
Authors not listed
This study presents a validation and refinement of the “yellow cards” error detection workflow that can be applied to any property connected to molecular structure. In our implementation the workflow employed 5 predictive models with each assigning a “yellow card” to 5% of the entries with worst prediction accuracy.…
Matthew W. Self, Peter Cheeseman
This paper shows that the common method used for making predictions under uncertainty in AI and science is in error. This method is to use currently available data to select the best model from a given class of models-this process is called abduction-and then to use this model to make predictions about future data. The…
Dylan Spicker, Amir Nazemi, Joy Hutchinson, Paul Fieguth + 3 more
Dietary intake data are routinely drawn upon to explore diet-health relationships, and inform clinical practice and public health. However, these data are almost always subject to measurement error, distorting true diet-health relationships. Beyond measurement error, there are likely complex synergistic and sometimes…
Hwiyoung Lee, Zhenyao Ye, Yun Yang, Yezhi Pan + 8 more
Machine learning (ML)- and artificial intelligence (AI)-based aging clocks are increasingly used to quantify physiological and molecular aging from omics and medical imaging data as distinct from chronological age. Here, we characterize a fundamental but underappreciated statistical limitation of commonly used ML/AI…
Hao Tang, Tianle Yue, Ying Li
Machine learning (ML) has become an important technique in materials science, markedly accelerating the discovery and design of novel materials, and concurrently lowering the burden of experimental costs. Uncertainty quantification (UQ) plays a pivotal role in the accurate prediction and innovative design of novel…
Danilo Bzdok, Denis Engemann, Olivier Grisel, Gaël Varoquaux + 1 more
In the 20^th^ century many advances in biological knowledge and evidence-based medicine were supported by p-values and accompanying methods. In the beginning 21^st^ century, ambitions towards precision medicine put a premium on detailed predictions for single individuals. The shift causes tension between traditional…
Esther Heid, Charles J. McGill, Florence H. Vermeire, William H. Green
Characterizing uncertainty in machine learning models has recently gained interest in the context of machine learning reliability, robustness, safety, and active learning. Here, we separate the total uncertainty into contributions from noise in the data (aleatoric) and shortcomings of the model (epistemic), further…
Authors not listed
Plastic mechanical recycling is the conventional technological step towards circularity. In such aspects, complex mixtures of polyolefin blends are often fed into mechanical recycling systems, resulting in moulded products with uncertain quality. To add to the difficulty of heterogeneous feedstocks, the testing of…
Authors not listed
The rapid growth of worldwide computing power has transformed in silico chemistry into a discipline that is integrated into the daily work of many chemists. Nowadays, researchers find it increasingly straightforward to predict a wide range of molecular properties and chemi- cal processes at reasonable computational…
Authors not listed
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) and Middle East respiratory syndrome coronavirus (MERS-CoV) are two important targets in current drug discovery, mainly due to the COVID-19 pandemic and the MERS-CoV outbreaks in recent years. An important target of both SARS-CoV-2 and MERS-CoV is the main…
Lu Hong, PJ Lamberson, Scott E Page
An increasing proportion of decisions, design choices, and predictions are being made by hybrid groups consisting of humans and artificial intelligence (AI). In this paper, we provide analytic foundations that explain the potential benefits of hybrid groups on predictive tasks, the primary use of AI. Our analysis…