23 papers · ranked by Valyu relevance
Shannon M. Locke, Michael S. Landy, Pascal Mamassian, Christoph Mathys
'Christoph Mathys'] Perceptual confidence is an important internal signal about the certainty of our decisions and there is a substantial debate on how it is computed. We highlight three confidence metric types from the literature: observers either use 1) the full probability distribution to compute probability correct…
Lorenzo Jaime Yu Flores, Ori Ernst, Jackie Kit Cheung
Well-calibrated model confidence scores can improve the usefulness of text generation models. For example, users can be prompted to review predictions with low confidence scores, to prevent models from returning bad or potentially dangerous predictions. However, confidence metrics are not always well calibrated in text…
Richard Oliver Lane
Probabilities or confidence values produced by artificial intelligence (AI) and machine learning (ML) models often do not reflect their true accuracy, with some models being under or over confident in their predictions. For example, if a model is 80% sure of an outcome, is it correct 80% of the time? Probability…
Pascal Mamassian, Vincent de Gardelle
Over the last decade, different approaches have been proposed to interpret confidence rating judgments obtained after perceptual decisions. One very popular approach is to compute meta-d’ which is a global measure of the sensibility to discriminate the confidence rating distributions for correct and incorrect…
Ziang Zhou, Tianyuan Jin, Jieming Shi, Qing Li
Large Language Models (LLMs) exhibit impressive performance across diverse domains but often suffer from overconfidence, limiting their reliability in critical applications. We propose SteerConf, a novel framework that systematically steers LLMs' confidence scores to improve their calibration and reliability. SteerConf…
Pascal Mamassian, Vincent de Gardelle, Christoph Strauch
Over the last decade, different approaches have been proposed to interpret confidence rating judgments obtained after perceptual decisions. One very popular approach is to compute meta-d’ which is a global measure of the sensibility to discriminate the confidence rating distributions for correct and incorrect…
Xiaoou Liu, Zhen Lin, Longchao Da, Chacha Chen + 2 more
Correctness Labels Authors: ['Xiaoou Liu' 'Zhen Lin' 'Longchao Da' 'Chacha Chen' 'Shubhendu Trivedi' 'Hua Quan Wei'] Large Language Models (LLMs) require robust confidence estimation, particularly in critical domains like healthcare and law where unreliable outputs can lead to significant consequences. Despite much…
Xiaojuan Ma, Xinru Wang, Ying Lei, Chuhan Shi + 2 more
'Xiaojuan Ma'] In AI-assisted decision-making, it is crucial but challenging for humans to achieve appropriate reliance on AI. This paper approaches this problem from a human-centered perspective, "human selfconfidence calibration". We begin by proposing an analytical framework to highlight the importance of calibrated…
Zoe M. Boundy-Singer, Corey M. Ziemba, Robbe L. T. Goris
Decisions vary in difficulty. Humans know this and typically report more confidence in easy than in difficult decisions. However, confidence reports do not perfectly track decision accuracy, but also reflect response biases and difficulty misjudgments. To isolate the quality of confidence reports, we developed a model…
Domenico Romanazzi, Anna Pasini, Luca Tarasi, Vincenzo Romei
Metacognition, the capacity to monitor and evaluate one’s own decisions, has become a central topic in psychological and neuroscientific research. While research has largely focused on metacognitive sensitivity, defined as the ability to discriminate between correct and incorrect decisions, considerably less attention…
Ji Won Bang, Medha Shekhar, Dobromir Rahnev
Visual metacognition is the ability to employ confidence ratings in order to predict the accuracy of ones decisions about visual stimuli. Despite years of research, it is still unclear how visual metacognitive efficiency can be manipulated. Here we show that a hierarchical model of confidence generation makes a…
Juhani Kivimäki, Jakub Białek, Wojtek Kuberski, Jukka K. Nurminen
Model monitoring is a critical component of the machine learning lifecycle, safeguarding against undetected drops in the model's performance after deployment. Traditionally, performance monitoring has required access to ground truth labels, which are not always readily available. This can result in unacceptable latency…
Houman Heidarabadi, Melina Graner, Holger Hesse
Profitability, reliability, and efficiency of battery systems across a broad spectrum of applications, including both stationary energy storage and automobile sectors, are critically dependent on accurate battery lifespan predic-tions. Traditional deterministic models for estimating battery longevity are inadequate, as…
Sucharit Katyal, Stephen M. Fleming
Foundational work in the psychology of metacognition identified a distinction between metacognitive knowledge (stable beliefs about one’s capacities) and metacognitive experiences (local evaluations of performance). More recently, the field has focused on developing tasks and metrics that seek to identify metacognitive…
Herrick Fung, N. Apurva Ratan Murty, Dobromir Rahnev
Human behavior differs substantially across individuals. While artificial neural networks (ANNs) are regarded as promising models of human perception, they are often assumed to lack such individual differences. Here, we demonstrate that multiple instances of the same ANN architecture exhibit substantial individual…
Timothy Gould
This study examined factors influencing student confidence and their perception of learning in the context of undergraduate chemistry and biochemistry courses. Anonymous online surveys were used to measure the extent to which small group work influenced student confidence in solving problems compared to working…
Benjamin Neely, Yasset Perez-Riverol, Magnus Palmblad
The past decade has seen widespread advances in quality control (QC) materials and software tools focused specifically on mass spectrometry-based proteomics, yet the rate of adoption is inconsistent. Despite the fundamental importance of QC, it typically falls behind learning new techniques, instruments, or software.…
Maria H. Rasmussen, Chenru Duan, Heather J. Kulik, Jan Halborg Jensen
With the increasingly more important role of machine learning (ML) models in chemical research, the need for putting a level of confidence to the model predictions naturally arises. Several methods for obtaining uncertainty estimates have been proposed in recent years but consensus on the evaluation of these have yet…
Erin C. Conrad, John M. Bernabei, Lohith G. Kini, Preya Shah + 6 more
Focal epilepsy is a clinical condition arising from disordered brain networks. Network models hold promise to map these networks, localize seizure generators, and inform targeted interventions to control seizures. However, incomplete sampling of epileptic brain due to sparse placement of intracranial electrodes may…
Robert Reischke
Confidence contours in parameter space are a helpful tool to compare and classify determined estimators. For more intricate parameter estimations of non-linear nature or complex error structures, the procedure of determining confidence contours is a statistically complex task. For polymer chemists, such particular…
Ningsheng Zhao, Trang Bui, Jia Yuan Yu, Krzysztof Dzieciolowski
Many classification performance metrics exist, each suited to a specific application. However, these metrics often differ in scale and can exhibit varying sensitivity to class imbalance rates in the test set. As a result, it is difficult to use the nominal values of these metrics to interpret and evaluate…
David Manheim
Title: Summary Metrics are useful for measuring systems and motivating behaviors in academia as well as in public policy, medicine, business, and other systems. Unfortunately, naive application of metrics to a system can distort the system and even undermine the original goal. There are two interrelated problems to…
Authors not listed
Accurately predicting the diverse bound-state conformations of small molecules is crucial for successful drug discovery and design, particularly when detailed protein-ligand interactions are unknown. Established tools exist, but efficiently exploring the vast conformational space remains challenging. This work…