22 papers · ranked by Valyu relevance
Gary S Collins, Paula Dhiman, Jie Ma, Michael M Schlussel + 9 more
Evaluating the performance of a clinical prediction model is crucial to establish its predictive accuracy in the populations and settings intended for use. In this article, the first in a three part series, Collins and colleagues describe the importance of a meaningful evaluation using internal, internal-external, and…
Ben Hutchinson, Negar Rostamzadeh, Christina M. Greer, Katherine Heller + 1 more
'Katherine Heller' 'Vinodkumar Prabhakaran'] Forming a reliable judgement of a machine learning (ML) model's appropriateness for an application ecosystem is critical for its responsible use, and requires considering a broad range of factors including harms, benefits, and responsibilities. In practice, however…
Utku Boran Torun, Veli Karakaya, Ali Babar, Eray Tüzün
Large Language Models (LLMs) are increasingly embedded in software engineering (SE) tools, powering applications such as code generation, automated code review, and bug triage. As these LLM-based AI for Software Engineering (AI4SE) systems transition from experimental prototypes to widely deployed tools, the question…
Maulik K. Nariya, Caitlin E. Mills, Peter K. Sorger, Artem Sokolov
The true accuracy of a machine learning model is a population-level statistic that cannot be observed directly. In practice, predictor performance is estimated against one or more test datasets, and the accuracy of this estimate strongly depends on how well the test sets represent all possible unseen datasets. Here we…
Kate E. Dray, Joseph J. Muldoon, Niall M. Mangan, Neda Bagheri + 1 more
Mathematical modeling is invaluable for advancing understanding and design of synthetic biological systems. However, the model development process is complicated and often unintuitive, requiring iteration on various computational tasks and comparisons with experimental data. Ad hoc model development can pose a barrier…
Pola Schwöbel, Luca Franceschi, Muhammad Bilal Zafar, Keerthan Vasist + 6 more
'Keerthan Vasist' 'A. Malhotra' 'Tomer Shenhar' 'Pinal Tailor' 'Pınar Yilmaz' 'Michael J. Diamond' 'Michele Donini'] fmeval is an open source library to evaluate large language models (LLMs) in a range of tasks. It helps practitioners evaluate their model for task performance and along multiple responsible AI…
Kenneth Benavides, Josh Fleischer, Danti Chen
Teams deploying large language models in business contexts need evaluation systems, yet most treat evaluation as static model selection: run benchmarks, rank models, deploy the winner. This framing misses evaluation's primary value for production systems--diagnosing why a system underperforms and guiding what to fix.…
Feng Feng, Zhenru Chen, Jianyuan Ni, Yuanxun Zhang + 3 more
Drinking water is essential to public health and socioeconomic growth. Therefore, assessing and ensuring drinking water supply is a critical task in modern society. Conventional approaches to analyzing and controlling drinking water quality are labor-intensive and costly with a low throughput. Machine learning (ML) is…
Thomas Wöhling, Alvaro Oliver Crespo Delgadillo, Moritz Kraft, Anneli Guthke
'Anneli Guthke'] Title: Abstract Groundwater level observations are used as decision variables for aquifer management, often in conjunction with models to provide predictions for operational forecasting. In this study, we compare different model classes for this task: a spatially explicit 3D groundwater flow model…
Keith D. Harris, Guy Hadari, Gili Greenbaum
Modelling the dynamics of biological processes is ubiquitous across the ecological and evolutionary disciplines. However, the increasing complexity of these models poses a challenge to the dissemination of model-derived results. Often only a small subset of model results are made available to the scientific community…
Birger Johansson, Trond A. Tjøstheim, Christian Balkenius
System-level brain modeling is a powerful method for building computational models of the brain and allows biologically motivated models to produce measurable behavior that can be tested against empirical data. System-level brain models occupy an intermediate position between detailed neuronal circuit models and…
Florian E. Dorner, Vivian Y. Nastl, Moritz Hardt
twice the data Authors: ['Florian E. Dorner' 'Vivian Y. Nastl' 'Moritz Hardt'] High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefore become an important research ambition. Many hope to use…
Henry Munroe, Bright Osatohanmwen, Reza Sharifi
Machine learning (ML) models with stochastic and non-deterministic characteristics are increasingly used for genomic prediction in plant breeding, but evaluation often neglects important aspects like prediction stability and ranking performance. This study addresses this gap by evaluating how two hyperparameters of a…
Houfang Guo, Muhammad Asif
Enterprises are urged to continue implementing the sustainable development strategy in their business operations as “carbon neutrality” and “carbon peak” gradually become the current stage’s worldwide targets. High-tech businesses (HTE) need to be better equipped to manage financial risks and avoid financial crises in…
Authors not listed
Background: Janus Kinase 2 (JAK2) is a key kinase in cellular signal transduction. Its abnormal activation is closely related to various myeloproliferative neoplasms and inflammatory diseases. Developing selective JAK2 inhibitors is an important direction in drug discovery. Accurate prediction of compound inhibitory…
Scott Spillias, Jacob Rogers, Fabio Boschetti, Beth Fulton + 3 more
Ecosystem models are essential for ecosystem management, but their development traditionally requires significant time and expertise, creating bottlenecks in addressing urgent environmental challenges. We present LEMMA (LLM Enabled Mechanistic Modelling for ecosystem Assessment), a framework that programmatically…
Fuzhan Rahmanian, Robert M. Lee, Dominik Linzner, Kathrin Michel + 4 more
Predicting and monitoring battery life early and across chemistries is a significant challenge due to the plethora of degradation paths, form factors, and electrochemical testing protocols. Existing models typically translate poorly across different electrode, electrolyte, and additive materials, mostly require a fixed…
Charlotte Christensen, André C. Ferreira, Wismer Cherono, Maria Maximiadi + 4 more
Machine-learning (ML) is revolutionizing the study of ecology and evolution, but the performance of models (and their evaluation) is dependent on the quality of the training and validation data. Currently, we have standard metrics for evaluating model performance (e.g., precision, recall, F1), but these to some extent…
Authors not listed
Machine learning holds significant promise for accelerating biomarker discovery in clinical proteomics, yet its real-world impact remains limited by widespread methodological pitfalls and unrealistic expectations. In this perspective, we critically examine the integration of machine learning into clinical proteomics…
David E Carlson, Ricardo Chavarriaga, Yiling Liu, Fabien Lotte + 1 more
'Bao-Liang Lu'] Title: Abstract Objective. Machine learning’s (MLs) ability to capture intricate patterns makes it vital in neural engineering research. With its increasing use, ensuring the validity and reproducibility of ML methods is critical. Unfortunately, this has not always been the case in practice, as there…
Daniel R Balcarcel, Sanjiv D Mehta, Celeste G Dixon, Charlotte Z Woods-Hill + 3 more
Prognostic models developed for use in the intensive care unit (ICU) can inform treatment decisions and improve patient care. However, despite extensive research, few models have contributed to improved patient-centred outcomes. A major limitation is that the influence of treatment interventions on patient outcomes…
Authors not listed
Computational blind challenges offer critical, unbiased assessment opportunities to assess and accelerate scientific progress, as demonstrated by a breadth of breakthroughs over the last decade. We report the outcomes and key insights from an open science community blind challenge focused on computational methods in…