24 papers · ranked by Valyu relevance
Randal S. Olson, William La Cava, Patryk Orzechowski, Ryan J. Urbanowicz + 1 more
'Ryan J. Urbanowicz' 'Jason H. Moore'] Background The selection, development, or comparison of machine learning methods in data mining can be a difficult task based on the target problem and goals of a particular study. Numerous publicly available real-world and simulated benchmark datasets have emerged from different…
Anasua Sarkar, Yang Yang, Mauno Vihinen
Development of new computational methods and testing their performance has to be done on experimental data. Only in comparison to existing knowledge can method performance be assessed. For that purpose, benchmark datasets with known and verified outcome are needed. High-quality benchmark datasets are valuable and may…
Bernd Bischl, Giuseppe Casalicchio, Matthias Feurer, Frank Hutter + 4 more
'Michel Lang' 'Rafael Gomes Mantovani' 'Jan N. van Rijn' 'Joaquin Vanschoren'] | Bernd Bischl | bernd.bischl@stat.uni-muenchen.de | | --- | --- | | Giuseppe Casalicchio | giuseppe.casalicchio@stat.uni-muenchen.de | | Matthias Feurer | feurerm@cs.uni-freiburg.de | | Frank Hutter | fh@cs.uni-freiburg.de | | Michel Lang |…
Jeyan Thiyagalingam, Mallikarjun Shankar, Geoffrey Fox, Tony Hey
The breakthrough in Deep Learning neural networks has transformed the use of AI and machine learning technologies for the analysis of very large experimental datasets. These datasets are typically generated by large-scale experimental facilities at national laboratories. In the context of science, scientific machine…
Roberta Rocca, Tal Yarkoni
Consensus on standards for evaluating models and theories is an integral part of every science. Nonetheless, in psychology, relatively little focus has been placed on defining reliable communal metrics to assess model performance. Evaluation practices are often idiosyncratic and are affected by a number of shortcomings…
Bo Wen, William Noble
Training machine learning models for tasks such as de novo sequencing or spectral clustering requires large collections of confidently identified spectra. Here we describe a dataset of 2.8 million high-confidence peptide-spectrum matches derived from nine different species. The dataset is based on a previously…
Sterling G. Baird, Taylor D. Sparks
In scientific disciplines, benchmarks play a vital role in driving progress forward. For a benchmark to be effective, it must closely resemble real-world tasks. If the level of difficulty or relevance is inadequate, it can impede progress in the field. Moreover, benchmarks should have low computational overhead to…
Kohulan Rajan, Henning Otto Brinkhaus, M. Isabel Agea, Achim Zielesny + 1 more
The number of publications describing chemical structures has increased steadily over the last decades. However, the majority of published chemical information is currently not available in machine-readable form in public databases. It remains a challenge to automate the process of information extraction in a way that…
Tristan Thrush, Kushal Tirumala, Anmol Gupta, Max Bartolo + 6 more
'Pedro Rodríguez' 'Tariq Kane' 'William Gaviria Rojas' 'Peter Mattson' 'Adina Williams' 'Douwe Kiela'] We introduce Dynatask: an open source system for setting up custom NLP tasks that aims to greatly lower the technical knowledge and effort required for hosting and evaluating state-of-the-art NLP models, as well as…
Zhen Xu, Sergio Escalera, Adrien Pavão, Magali Richard + 4 more
'Quanming Yao' 'Huan Zhao' 'Isabelle Guyon'] Title: Summary Obtaining a standardized benchmark of computational methods is a major issue in data-science communities. Dedicated frameworks enabling fair benchmarking in a unified environment are yet to be developed. Here, we introduce Codabench, a meta-benchmark platform…
Bernd Bischl, Giuseppe Casalicchio, Taniya Das, Matthias Feurer + 13 more
'Sebastian Fischer' 'Pieter Gijsbers' 'Subhaditya Mukherjee' 'Andreas C. Müller' 'László Németh' 'Luis Oala' 'Lennart Purucker' 'Sahithya Ravi' 'Jan N. van Rijn' 'Prabhant Singh' 'Joaquin Vanschoren' 'Jos van der Velde' 'Marcel Wever'] Title: Summary OpenML is an open-source platform that democratizes machine-learning…
Xiaoqi Cabiria Liang, Nick Robertson, Marni Torkel, Sanghyun Kim + 3 more
The rapid growth of computational methods for the computational biology field highlights the critical role of benchmarking in guiding method selection. However, there is no standardised data structure that effectively links and stores datasets, performance metrics and available ground truth. Without such a unified and…
Izaskun Mallona, Charlotte Soneson, Ben Carrillo, Almut Lütge + 5 more
Benchmarking, which involves collecting reference datasets and demonstrating method performance, is a requirement for the development of new computational tools, but also becomes a domain of its own to achieve neutral comparisons of methods. Although a lot has been written about how to design and conduct benchmark…
Avanika Narayan, Piero Molino, Karan Goel, Willie Neiswanger + 1 more
'Christopher Ré'] The rapid proliferation of machine learning models across domains and deployment settings has given rise to various communities (e.g. industry practitioners) which seek to benchmark models across tasks and objectives of personal value. Unfortunately, these users cannot use standard benchmark results…
Zhi Xu, Sérgio Escalera, Adrien Pavao, Magali Richard + 4 more
'Quanming Yao' 'Huan Zhao' 'Isabelle Guyon'] - We developed Codabench, an open-source and community-driven benchmark platform, which facilitates benchmarking and guarantees reproducibility. - The platform allows organizers to set up benchmarks with custom designs and welcome contributions of users in the form of…
Authors not listed
The incredible advances in deep learning architectures have not only led to successful protein structure prediction methods, but also to many interesting new methods related to protein-ligand docking and co-folding. The most recent biomolecular foundation model, Boltz-2, even claimed to approach the performance of…
Salvador Capella-Gutierrez, Diana de la Iglesia, Juergen Haas, Analia Lourenco + 7 more
The dependence of life scientists on software has steadily grown in recent years. For many tasks, researchers have to decide which of the available bioinformatics software are more suitable for their specific needs. Additionally researchers should be able to objectively select the software that provides the highest…
Authors not listed
The construction of large benchmark sets has accelerated advancement of quantum chemistry methods, especially in density functional theory and lower-cost methods. However, these large benchmark sets can be unsuitable for cutting-edge method development, because research codes developed for fundamentally new approaches…
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton + 1 more
'Emily Denton' 'Alex Hanna'] Title: Summary In this work, we survey a breadth of literature that has revealed the limitations of predominant practices for dataset collection and use in the field of machine learning. We cover studies that critically review the design and development of datasets with a focus on negative…
Ulysse Rubens, Romain Mormont, Lassi Paavolainen, Volker Bäcker + 15 more
Automated image analysis has become key to extract quantitative information from scientific microscopy bioimages, but the methods involved are now often so refined that they can no longer be unambiguously described using written protocols. We introduce BIAFLOWS, a software tool with web services and a user interface…
Riley Hickman, Priyansh Parakh, Austin Cheng, Qianxiang Ai + 3 more
Experiment planning algorithms are a required component of autonomous platforms for scientific discovery. Selecting a suitable optimization algorithm for a novel application is an important yet difficult choice a researcher has to make based on past empirical performance on similar tasks. To facilitate the evaluation…
Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao + 4 more
'Neil Houlsby' 'Fernando Díaz' 'Donald Metzler' 'Oriol Vinyals'] The world of empirical machine learning (ML) strongly relies on benchmarks in order to determine the relative effectiveness of different algorithms and methods. This paper proposes the notion of a benchmark lottery that describes the overall fragility of…
Anthony Sonrel, Almut Luetge, Charlotte Soneson, Izaskun Mallona + 16 more
Computational methods represent the lifeblood of modern molecular biology. Benchmarking is important for all methods, but with a focus here on computational methods, benchmarking is critical to dissect important steps of analysis pipelines, formally assess performance across common situations as well as edge cases, and…
Pavan Kumar Behara, Hyesu Jang, Joshua Horton, Trevor Gokey + 6 more
A wide range of density functional methods and basis sets are available to derive the electronic structure and properties of molecules. Quantum mechanical calculations are too computationally intensive for routine simulation of molecules in the condensed phase, prompting the development of computationally efficient…