25 papers · ranked by Valyu relevance
Roberta Rocca, Tal Yarkoni
Consensus on standards for evaluating models and theories is an integral part of every science. Nonetheless, in psychology, relatively little focus has been placed on defining reliable communal metrics to assess model performance. Evaluation practices are often idiosyncratic and are affected by a number of shortcomings…
Gábor Szárnyas, Benedek Izsó, István Ráth, Dániel Varró
In model-driven development of safety-critical systems (like automotive, avionics or railways), well-formedness of models is repeatedly validated in order to detect design flaws as early as possible. In many industrial tools, validation rules are still often implemented by a large amount of imperative model traversal…
Jianfeng Zhan, Lei Wang, Wanling Gao, Hongxiao Li + 9 more
'Yunyou Huang' 'Yatao Li' 'Zhengxin Yang' 'Guoxin Kang' 'Chunjie Luo' 'Hainan Ye' 'Shaopeng Dai' 'Zhifei Zhang'] Evaluation is a crucial aspect of human existence and plays a vital role in various fields. However, it is often approached in an empirical and ad-hoc manner, lacking consensus on universal concepts…
Salvador Capella-Gutierrez, Diana de la Iglesia, Juergen Haas, Analia Lourenco + 7 more
The dependence of life scientists on software has steadily grown in recent years. For many tasks, researchers have to decide which of the available bioinformatics software are more suitable for their specific needs. Additionally researchers should be able to objectively select the software that provides the highest…
Zhen Xu, Sergio Escalera, Adrien Pavão, Magali Richard + 4 more
'Quanming Yao' 'Huan Zhao' 'Isabelle Guyon'] Title: Summary Obtaining a standardized benchmark of computational methods is a major issue in data-science communities. Dedicated frameworks enabling fair benchmarking in a unified environment are yet to be developed. Here, we introduce Codabench, a meta-benchmark platform…
Jin Liu, Qingquan Li, Wenlong Du
Large Language Models Authors: ['Jin Liu' 'Qingquan Li' 'Wenlong Du'] In current benchmarks for evaluating large language models (LLMs), there are issues such as evaluation content restriction, untimely updates, and lack of optimization guidance. In this paper, we propose a new paradigm for the measurement of LLMs…
Xiaoqi Cabiria Liang, Nick Robertson, Marni Torkel, Sanghyun Kim + 3 more
The rapid growth of computational methods for the computational biology field highlights the critical role of benchmarking in guiding method selection. However, there is no standardised data structure that effectively links and stores datasets, performance metrics and available ground truth. Without such a unified and…
Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning + 2 more
Despite substantial excitement around the use of AI in law, little information exists on the performance and associated risks of the domain’s widely marketed tools. Recent work, for instance, has demonstrated the significant potential for “hallucinations”-wherein models make up facts, law, and precedent-leading Chief…
Martin Grambow, Christoph Laaber, Philipp Leitner, David Bermbach + 1 more
'Muhammad Aleem'] Performance problems in applications should ideally be detected as soon as they occur, i.e., directly when the causing code modification is added to the code repository. To this end, complex and cost-intensive application benchmarks or lightweight but less relevant microbenchmarks can be added to…
Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman + 14 more
The rapid scaling of Genomic Foundation Models (GFMs) has created a critical need for standardized evaluation frameworks. Current benchmarking practices are often fragmented, relying on model-specific preprocessing and inconsistent metric implementations that hinder reproducible comparisons. We present GFMBench-API, a…
Anthony Sonrel, Almut Luetge, Charlotte Soneson, Izaskun Mallona + 16 more
Computational methods represent the lifeblood of modern molecular biology. Benchmarking is important for all methods, but with a focus here on computational methods, benchmarking is critical to dissect important steps of analysis pipelines, formally assess performance across common situations as well as edge cases, and…
Marc Pagès-Gallego, Jeroen de Ridder
Nanopore basecalling is a difficult task which requires the use of complex algorithms and neural network models to achieve competitive accuracies. These algorithms, called basecallers, are developed continuously both by ONT and the scientific community in an effort to improve basecalling accuracies. With the rapidly…
Ian Knight, Khanh Tang, Olivier Mailhot, John Irwin
Molecular docking is a widely used technique for leveraging protein structure in ligand discovery, but as a method, it remains difficult to utilize due to limitations that have not been adequately addressed. Despite some progress towards automation, docking still requires expert guidance, hindering its adoption by a…
Zheng Li, Liam O’Brien, Maria Kihl
Software engineering considers performance evaluation to be one of the key portions of software quality assurance. Unfortunately, there seems to be a lack of standard methodologies for performance evaluation even in the scope of experimental computer science. Inspired by the concept of "instantiation" in…
A. Wind, W. H. van Harten
Background Although benchmarking may improve hospital processes, research on this subject is limited. The aim of this study was to provide an overview of publications on benchmarking in specialty hospitals and a description of study characteristics. Methods We searched PubMed and EMBASE for articles published in…
Izaskun Mallona, Almut Luetge, Ben Carrillo, Daniel Incicau + 4 more
bioinformatics Authors: ['Izaskun Mallona' 'Almut Luetge' 'Ben Carrillo' 'Daniel Incicau' 'Reto Gerber' 'Anthony Sonrel' 'Charlotte Soneson' 'Mark D. Robinson'] We describe an alpha version of a new benchmarking system, Omnibenchmark, to facilitate benchmark formalization and execution in solo and community efforts.…
Authors not listed
Accurate benchmarks are key to assessing the accuracy and robustness of computational methods, yet most available benchmark sets focus on equilibrium geometries, limiting their utility for applications involving non-equilibrium structures such as ab initio molecular dynamics and automated reaction-path exploration. To…
Jianfeng Zhan
Currently, there is no consistent benchmarking across multi-disciplines. Even no previous work tries to relate different categories of benchmarks in multi-disciplines. This article investigates the origin and evolution of the benchmark term. Five categories of benchmarks are summarized, including measurement standards…
Morgan Thomas, Noel M. O'Boyle, Andreas Bender, Chris de Graaf
MolScore is an open-source Python framework for scoring and evaluating molecules in the context of goal-directed generative models as used in de novo drug design. MolScore includes many relevant scoring functions for de novo drug design such as molecular similarity, docking software, predictive models, and…
Riley Hickman, Priyansh Parakh, Austin Cheng, Qianxiang Ai + 3 more
Experiment planning algorithms are a required component of autonomous platforms for scientific discovery. Selecting a suitable optimization algorithm for a novel application is an important yet difficult choice a researcher has to make based on past empirical performance on similar tasks. To facilitate the evaluation…
Authors not listed
Proteochemometric models (PCM) are used in computational drug discovery to leverage both protein and ligand representations for bioactivity prediction. While machine learning (ML) and deep learning (DL) have come to dominate PCMs, often serving as scoring functions, rigorous evaluation standards have not always been…
Roham Koohestani, Philippe de Bekker, Maliheh Izadi
—Benchmarks are essential for consistent evaluation and reproducibility. The integration of Artificial Intelligence into Software Engineering (AI4SE) has given rise to numerous benchmarks for tasks such as code generation and bug fixing. However, this surge presents challenges: (1) scattered benchmark knowledge across…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Sterling G. Baird, Taylor D. Sparks
In scientific disciplines, benchmarks play a vital role in driving progress forward. For a benchmark to be effective, it must closely resemble real-world tasks. If the level of difficulty or relevance is inadequate, it can impede progress in the field. Moreover, benchmarks should have low computational overhead to…
Yu-Chieh Huang, Pierre Tremouilhac, Stefan Kuhn, Pei-Chi Huang + 6 more
A method for data review in chemical sciences with a focus on data for the characterization of synthetic molecules is described. As current procedures for data curation in chemistry rely almost exclusively on manual checking or peer reviewing, a (semi-)automatic procedure for the evaluation of data assigned to…