22 papers · ranked by Valyu relevance
Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec, Jan Brauner + 1 more
'Matthias Samwald'] Benchmarks are crucial to measuring and steering progress in artificial intelligence (AI). However, recent studies raised concerns over the state of AI benchmarking, reporting issues such as benchmark overfitting, benchmark saturation and increasing centralization of benchmark dataset creation. To…
Burden, John, Tešić, Marko + 4 more
Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation paradigms have emerged, often developing in isolation, adopting conflicting terminologies, and overlooking each other's contributions. This…
Zerui Cheng, Stella Wohnig, Ruchika Gupta, Samiul Alam + 13 more
Zerui Cheng1, ∗ Stella Wohnig2, ∗ Ruchika Gupta3, ∗ Samiul Alam4, ∗ Tassallah Abdullahi 5 João Alves Ribeiro 6 Christian Nielsen-Garcia 7 Saif Mir 4 Siran Li 8 Jason Orender 9 Seyed Ali Bahrainian 8 Daniel Kirste 10 Aaron Gokaslan 11 Mikołaj Glinka 12 Carsten Eickhoff8,† Ruben Wolff12,† 1 Princeton University 2 CISPA…
von Laszewski, Gregor, Brewer, Wesley + 58 more
| I | | Introduction | 3 | |-----|-----------------|--------------------------------------------------------------------------------|----| | II | Definitions | | 4 | | | II-A | What is Benchmarking? | 4 | | | II-B | Lessons Learned from Traditional HPC Benchmarking | 4 | | | II-C | What is Democratization? | 4 | | | |…
Maria Eriksson, Erasmo Purificato, A. Noroozian, João Vinagre + 3 more
Issues in AI Evaluation Authors: ['Maria Eriksson' 'Erasmo Purificato' 'A. Noroozian' 'João Vinagre' 'Guillaume Chaslot' 'Emilia Gómez' 'David Fernández Llorca'] MARIA ERIKSSON∗ , European Commission, Joint Research Centre (JRC), Seville, Spain ERASMO PURIFICATO∗ , European Commission, Joint Research Centre (JRC)…
Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton + 1 more
'Emily Denton' 'Alex Hanna'] There is a tendency across different subfields in AI to valorize a small collection of influential benchmarks. These benchmarks operate as stand-ins for a range of anointed common problems that are frequently framed as foundational milestones on the path towards flexible and generalizable…
Jacqueline Michelle Metsch, Anne-Christin Hauschild
The increasing digitalisation of multi-modal data in medicine and novel artificial intelligence (AI) algorithms opens up a large number of opportunities for predictive models. In particular, deep learning models show great performance in the medical field. A major limitation of such powerful but complex models…
Zhen Xu, Sergio Escalera, Adrien Pavão, Magali Richard + 4 more
'Quanming Yao' 'Huan Zhao' 'Isabelle Guyon'] Title: Summary Obtaining a standardized benchmark of computational methods is a major issue in data-science communities. Dedicated frameworks enabling fair benchmarking in a unified environment are yet to be developed. Here, we introduce Codabench, a meta-benchmark platform…
Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning + 2 more
Despite substantial excitement around the use of AI in law, little information exists on the performance and associated risks of the domain’s widely marketed tools. Recent work, for instance, has demonstrated the significant potential for “hallucinations”-wherein models make up facts, law, and precedent-leading Chief…
Alexander Campolo
This commentary situates the epistemic values of machine learning’s culture of benchmarking and evaluation within larger temporal structures. Beyond questions of validity, whether model comparisons are statistically valid or whether benchmarks adequately represent meaningful tasks or capabilities, it asks how…
Stefan Baack, Christo Buschek, Maty Bohacek
The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases and company blog posts, where model builders highlight results on selected benchmarks. These artifacts now largely define the state of the art for researchers and the…
Travis LaCroix, Alexandra Sasha Luccioni
Benchmarks are seen as the cornerstone for measuring technical progress in Artificial Intelligence (AI) research and have been developed for a variety of tasks ranging from question answering to facial recognition. An increasingly prominent research area in AI is ethics, which currently has no set of benchmarks nor…
Cody Kommers, Ruth Ahnert, Maria Antoniak, Emmanouil Benetos + 34 more
Generative AI (GenAI) systems are increasingly recognized as cultural technologies, yet current evaluation frameworks often treat culture as a variable to be measured rather than fundamental to the system's operation. Drawing on hermeneutic theory from the humanities, we argue that GenAI systems function as "context…
Wenbin Guo, Minzhe Zhang, Bowei Han, Youjia Ma + 6 more
Large language model (LLM)-based agents hold transformative potential for automating bioinformatics workflows; however, systematic evaluations of their capabilities remain limited, hindering a clear assessment of their readiness for real-world application. We introduce PromptBio-Bench, a comprehensive evaluation suite…
Henry E. Miller, Matthew Greenig, Benjamin Tenmann, Bo Wang
Large language model (LLM) agents hold promise for accelerating biomedical research and development (R&D). Several biomedical agents have recently been proposed, but their evaluation has largely been restricted to question answering (e.g., LAB-Bench) or narrow bioinformatics tasks. Presently, there remains a lack of…
Eyes S. Robson, Nilah M. Ioannidis
Computational genomics increasingly relies on machine learning methods for genome interpretation, and the recent adoption of neural sequence-to-function models highlights the need for rigorous model specification and controlled evaluation, problems familiar to other fields of AI. Research strategies that have greatly…
Authors not listed
Accurate benchmarks are key to assessing the accuracy and robustness of computational methods, yet most available benchmark sets focus on equilibrium geometries, limiting their utility for applications involving non-equilibrium structures such as ab initio molecular dynamics and automated reaction-path exploration. To…
Authors not listed
The AIQM series methods are successful neural network-based models that target coupled-cluster accuracy while maintaining high robustness and transferability across various tasks by leveraging Δ-learning. However, the previous AIQM1 and AIQM2 models are limited to molecular systems with four elements: H, C, N, and O…
Xiaoqi Cabiria Liang, Nick Robertson, Marni Torkel, Sanghyun Kim + 3 more
The rapid growth of computational methods for the computational biology field highlights the critical role of benchmarking in guiding method selection. However, there is no standardised data structure that effectively links and stores datasets, performance metrics and available ground truth. Without such a unified and…
Henning Otto Brinkhaus, Kohulan Rajan, Jonas Schaub, Achim Zielesny + 1 more
Recent years have seen a sharp increase in the development of deep learning and artificial intelligence-based molecular informatics. There has been a growing interest in applying deep learning to several subfields, including the digital transformation of synthetic chemistry, extraction of chemical information from the…
Authors not listed
Artificial intelligence (AI) is poised to transform heterogeneous catalysis, ushering in a new paradigm for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental, and chemical…
Authors not listed
This article proposes a three-level classification of artificial intelligence (AI) application in chemical sciences, reflecting the increasing degree of technology involvement in scientific and production processes: from automation of routine tasks (the level of "AI Assistant"), to the creation of specialized…