13 papers · ranked by Valyu relevance
Randal S. Olson, William La Cava, Patryk Orzechowski, Ryan J. Urbanowicz + 1 more
'Ryan J. Urbanowicz' 'Jason H. Moore'] Background The selection, development, or comparison of machine learning methods in data mining can be a difficult task based on the target problem and goals of a particular study. Numerous publicly available real-world and simulated benchmark datasets have emerged from different…
Han Xing, Yongjian Zhao
Predicting drug-diseases associations provides hints in developing new drugs. Various computational methods have been developed. To develop better models for predicting drug-disease associations, two types of key resources must be obtained, the benchmarking dataset and the baseline performances. Collecting these…
Trishan Panch, Tom J. Pollard, Heather Mattie, Emily Lindemer + 2 more
Benchmark datasets have a powerful normative influence: by determining how the real world is represented in data, they define which problems will first be solved by algorithms built using the datasets and, by extension, who these algorithms will work for. It is desirable for these datasets to serve four functions: (1)…
Roberta Rocca, Tal Yarkoni
Consensus on standards for evaluating models and theories is an integral part of every science. Nonetheless, in psychology, relatively little focus has been placed on defining reliable communal metrics to assess model performance. Evaluation practices are often idiosyncratic and are affected by a number of shortcomings…
Zhen Xu, Sergio Escalera, Adrien Pavão, Magali Richard + 4 more
'Quanming Yao' 'Huan Zhao' 'Isabelle Guyon'] Title: Summary Obtaining a standardized benchmark of computational methods is a major issue in data-science communities. Dedicated frameworks enabling fair benchmarking in a unified environment are yet to be developed. Here, we introduce Codabench, a meta-benchmark platform…
Alexandros Karargyris, Renato Umeton, Micah J. Sheller, Alejandro Aristizabal + 70 more
Medical artificial intelligence (AI) has tremendous potential to advance healthcare by supporting and contributing to the evidence-based practice of medicine, personalizing patient treatment, reducing costs, and improving both healthcare provider and patient experience. Unlocking this potential requires systematic…
Serghei Mangul, Lana S. Martin, Brian L. Hill, Angela Ka-Mei Lam + 4 more
'Margaret G. Distler' 'Alex Zelikovsky' 'Eleazar Eskin' 'Jonathan Flint'] Computational omics methods packaged as software have become essential to modern biological research. The increasing dependence of scientists on these powerful software tools creates a need for systematic assessment of these methods, known as…
Bernd Bischl, Giuseppe Casalicchio, Taniya Das, Matthias Feurer + 13 more
'Sebastian Fischer' 'Pieter Gijsbers' 'Subhaditya Mukherjee' 'Andreas C. Müller' 'László Németh' 'Luis Oala' 'Lennart Purucker' 'Sahithya Ravi' 'Jan N. van Rijn' 'Prabhant Singh' 'Joaquin Vanschoren' 'Jos van der Velde' 'Marcel Wever'] Title: Summary OpenML is an open-source platform that democratizes machine-learning…
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton + 1 more
'Emily Denton' 'Alex Hanna'] Title: Summary In this work, we survey a breadth of literature that has revealed the limitations of predominant practices for dataset collection and use in the field of machine learning. We cover studies that critically review the design and development of datasets with a focus on negative…
Izaskun Mallona, Charlotte Soneson, Ben Carrillo, Almut Lütge + 5 more
Benchmarking, which involves collecting reference datasets and demonstrating method performance, is a requirement for the development of new computational tools, but also becomes a domain of its own to achieve neutral comparisons of methods. Although a lot has been written about how to design and conduct benchmark…
Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec, Jan Brauner + 1 more
'Matthias Samwald'] Benchmarks are crucial to measuring and steering progress in artificial intelligence (AI). However, recent studies raised concerns over the state of AI benchmarking, reporting issues such as benchmark overfitting, benchmark saturation and increasing centralization of benchmark dataset creation. To…
Kyle Ellrott, Alex Buchanan, Allison Creason, Michael Mason + 10 more
'Thomas Schaffter' 'Bruce Hoff' 'James Eddy' 'John M. Chilton' 'Thomas Yu' 'Joshua M. Stuart' 'Julio Saez-Rodriguez' 'Gustavo Stolovitzky' 'Paul C. Boutros' 'Justin Guinney'] Challenges are achieving broad acceptance for addressing many biomedical questions and enabling tool assessment. But ensuring that the methods…
Gábor Szárnyas, Benedek Izsó, István Ráth, Dániel Varró
In model-driven development of safety-critical systems (like automotive, avionics or railways), well-formedness of models is repeatedly validated in order to detect design flaws as early as possible. In many industrial tools, validation rules are still often implemented by a large amount of imperative model traversal…