21 papers · ranked by Valyu relevance
Janet, Lin, Zhang, Liangwei
This chapter bridges technical analysis and organizational preparedness by tracing the path from layered failure modes to reliability awareness in generative and agentic AI systems. We first introduce an 11-layer failure stack, a structured framework for identifying vulnerabilities ranging from hardware and power…
Yili Hong, Jiayi Lian, Li Xu, Jie Min + 3 more
'Xinwei Deng'] Artificial intelligence (AI) systems have become increasingly popular in many areas. Nevertheless, AI technologies are still in their developing stages, and many issues need to be addressed. Among those, the reliability of AI systems needs to be demonstrated so that the AI systems can be used with…
Simin Zheng, Jeanne M. Clark, Fatemeh Salboukh, Priscila Silva + 9 more
DR-AIR, a Comprehensive Data Repository for AI Reliability Authors: ['Simin Zheng' 'Jeanne M. Clark' 'Fatemeh Salboukh' 'Priscila Silva' 'Karen da Mata' 'Fenglian Pan' 'Jie Min' 'Jiayi Lian' 'Caleb King' 'Lance Fiondella' 'Jian Liu' 'Xinwei Deng' 'Yili Hong'] Artificial intelligence (AI) technology and systems have…
Authors not listed
AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in practice. This discrepancy highlights a major limitation of current evaluations: focusing on a single metric is not enough to understand agent…
Andrea Ferrario
We address an open problem in the philosophy of artificial intelligence (AI): how to justify the epistemic attitudes we have towards the trustworthiness of AI systems. The problem is important, as providing reasons to believe that AI systems are worthy of trust is key to appropriately rely on these systems in human-AI…
Juan M. Durán, David van der Vloed, Arnout Ruifrok, Rolf J.F. Ypma
Techniques from artificial intelligence (AI) can be used in forensic evidence evaluation and are currently applied in biometric fields. However, it is generally not possible to fully understand how and why these algorithms reach their conclusions. Whether and how we should include such ‘black box’ algorithms in this…
Luisa Roeder, Pamela Hoyte, Johan van der Meer, Lauren Fell + 4 more
'Patrick Johnston' 'Graham Kerr' 'Peter Bruza' 'Irina Basieva'] This exploratory study investigates a human agent’s evolving judgements of reliability when interacting with an AI system. Two aims drove this investigation: (1) compare the predictive performance of quantum vs. Markov random walk models regarding human…
Sabryn Hamila, Kyle Birchill, Khoa Cao, Md Nassif Hossain + 5 more
Background Multiple mini interviews (MMIs) are widely used in medical school admissions to assess applicants’ nonacademic attributes in a structured and reliable manner. However, the development of high-quality MMI stations is resource intensive and dependent on expert input. Objective This study explored the utility…
Adam M. Ostrovsky, Joshua R. Chen, Vishal N. Shah, Babak Abai
To the Editor: We appreciate the thoughtful and comprehensive response to our paper, Performance of 5 prominent large language models in surgical knowledge evaluation: a comparative analysis.1 We are grateful for the opportunity to address the important points raised by the authors of the letter.2 We concur with the…
John Dorsch, Ophelia Deroy
This study explores whether labeling AI as either “trustworthy” or “reliable” influences user perceptions and acceptance of automotive AI technologies. Utilizing a one-way between-subjects design, the research presented online participants (N = 478) with a text presenting guidelines for either trustworthy or reliable…
Amber W. Lockrow, Roni Setton, Karen A.P. Spreng, Signy Sheldon + 2 more
Autobiographical memory (AM) involves a rich phenomenological re-experiencing of a spatio-temporal event from the past, which is challenging to objectively quantify. The Autobiographical Interview (AI; Levine et al., 2002, Psychology & Aging) is a manualized performance-based assessment designed to quantify episodic…
Mats Tveter, Thomas Tveitstøl, Christoffer Hatlestad-Hall, Hugo L. Hammer + 1 more
As artificial intelligence (AI) is increasingly integrated into medical diagnostics, it is essential that predictive models provide not only accurate outputs but also reliable estimates of uncertainty. In clinical applications, where decisions have significant consequences, understanding the confidence behind each…
Authors not listed
The value of generative artificial intelligence (AI) for teaching and learning is currently hotly debated. Concerns regarding the accuracy of information produced by generative AI as well as student over-reliance on this tool coexist with excitement about tailored opportunities that AI may provide for educational…
Granville J. Matheson
Positron emission tomography (PET), along with many other fields of clinical research, is both timeconsuming and expensive, and recruitable patients can be scarce. These constraints limit the possibility of large-sample experimental designs, and often lead to statistically underpowered studies. This problem is…
Authors not listed
The description of heterogeneous catalysis is challenged by the intricacy of numerous multi-scale processes that govern the performance of catalyst materials. The chemical environment of the catalytic process and the kinetics of structural changes create configurations of typically unknown local geometries and…
Authors not listed
As the utilization of artificial intelligence (AI) and generative AI (GenAI) is expanding in the educational field, presenting significant implications for STEM disciplines, it is bringing opportunities to enhance how chemistry and chemical engineering are taught and learned. This perspective critically explores the…
Luisa Roeder, Pamela Hoyte, Graham Kerr, Peter Bruza + 1 more
This study provides an integrated electrophysiological and behavioral account of the neuro-cognitive markers underlying trust evolution during human interaction with artificial intelligence (AI). Trust is essential for effective collaboration and plays a key role in realizing the benefits of human–AI teaming in…
Authors not listed
Machine learning (ML) models are increasingly used in quantum chemistry, but their reliability hinges on uncertainty quantification (UQ). In this study, we compare two prominent UQ paradigms—Deep Evidential Regression (DER) and Deep Ensembles—on the QM9 and WS22 datasets, with a specific emphasis on the role of post…
Authors not listed
Artificial intelligence (AI) is reshaping chemical engineering. Still, its role in safety-critical operations is limited because we rarely see tools that link physical models with data-driven methods. This study brings together three elements: physics-constrained neural networks, uncertainty quantification, and a…
Authors not listed
Artificial intelligence (AI) is reshaping scientific research by accelerating discovery and enabling the analysis of complex data that traditional methods struggle to handle. This review examines over 310,000 journal articles and patents from the CAS Content Collection (2015–2025), with a focus on, biomedical research…
Authors not listed
This article proposes a three-level classification of artificial intelligence (AI) application in chemical sciences, reflecting the increasing degree of technology involvement in scientific and production processes: from automation of routine tasks (the level of "AI Assistant"), to the creation of specialized…