24 papers · ranked by Valyu relevance
von Laszewski, Gregor, Brewer, Wesley + 58 more
| I | | Introduction | 3 | |-----|-----------------|--------------------------------------------------------------------------------|----| | II | Definitions | | 4 | | | II-A | What is Benchmarking? | 4 | | | II-B | Lessons Learned from Traditional HPC Benchmarking | 4 | | | II-C | What is Democratization? | 4 | | | |…
Alexander Partin, Priyanka Vasanthakumari, Oleksandr Narykov, Andreas Wilke + 16 more
Deep learning and machine learning models have shown promise in drug response prediction (DRP), yet their ability to generalize across datasets remains an open question, raising concerns about their real-world applicability. Due to the lack of standardized benchmarking approaches, model evaluations and comparisons…
Philipp D. Siedler, Jordan Sassoon
Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. We introduce a dataset-centric meta-evaluation framework that audits benchmark datasets at the sample level along five…
Authors not listed
Accurately measuring compound binding affinities is key to driving the pharmaceutical development process. Rigorous physics-based in silico approaches, particularly alchemical free energy methods, have become a gold standard tool for estimating compound affinity changes. Here we present the results of a large-scale…
Authors not listed
Target-aware molecular generation models have emerged as promising tools for structure-based drug discovery, yet it remains unclear whether they genuinely exploit target information or merely resemble the Texas Sharpshooter fallacy by retrospectively rationalizing outputs. To address this, we introduce TarPass, a…
Lulu Wang, Zhuhuang Zhou
Background: Artificial intelligence (AI) has shown promise in breast ultrasound image analysis, but most evidence still comes from single-dataset studies. Clinical translation requires evaluation under heterogeneous acquisition and curation conditions. This study presents a patient-leakage-aware, reproducible benchmark…
Authors not listed
We report a new charge model and a new general small molecule force field. Here, we address the development and benchmarking of both the Open Force Field (OpenFF) AshGC charge model, as well as the Sage 2.3.0 small molecule force field for drug-like molecules. AshGC is a graph neural network-based method for efficient…
Authors not listed
The AIQM series methods are successful neural network-based models that target coupled-cluster accuracy while maintaining high robustness and transferability across various tasks by leveraging Δ-learning. However, the previous AIQM1 and AIQM2 models are limited to molecular systems with four elements: H, C, N, and O…
Mohamed M. Abbassy, Waleed M. Ead, Amr I. A. El-Shora, Ayman M. Aboalndr
In the rapidly evolving landscape of Natural Language Processing (NLP), transfer learning has emerged as a game-changing methodology, fundamentally altering how machine learning models are trained and deployed. The study at hand dives deep into the intricacies of transfer learning techniques, specifically focusing on…
Zerui Cheng, Stella Wohnig, Ruchika Gupta, Samiul Alam + 13 more
Zerui Cheng1, ∗ Stella Wohnig2, ∗ Ruchika Gupta3, ∗ Samiul Alam4, ∗ Tassallah Abdullahi 5 João Alves Ribeiro 6 Christian Nielsen-Garcia 7 Saif Mir 4 Siran Li 8 Jason Orender 9 Seyed Ali Bahrainian 8 Daniel Kirste 10 Aaron Gokaslan 11 Mikołaj Glinka 12 Carsten Eickhoff8,† Ruben Wolff12,† 1 Princeton University 2 CISPA…
Xuehua Zhou, Hanming Zhang, Tiantian Du, Quanbo Yuan + 2 more
Accurate prediction of marine and atmospheric environmental variables is important for climate adaptation, ecosystem management, and operational decision-making, yet practitioners still lack clear guidance on which machine-learning models are reliable across heterogeneous environmental tasks. We therefore developed a…
Authors not listed
Recent years have seen a growing interest in machine learning approaches for chemical tasks. The best existing methods focus on building base models that combine molecular graphs (“2D structures”) with atomic coordinates in 3D to predict molecular properties, typically through pre-training followed by fine-tuning on…
Hawks, Ben, von Laszewski, Gregor + 14 more
—Scientific machine learning research spans diverse domains and data modalities, yet existing benchmark efforts remain siloed and lack standardization. This makes novel and transformative applications of machine learning to critical scientific use-cases more fragmented and less clear in pathways to impact. This paper…
Robbe Devreese, Caroline Jachmann, Bart Van Puyvelde, Holda A. Anagho-Mattanovich + 46 more
Mass spectrometry (MS)-based proteomics is a well-established strategy for analyzing complex biological mixtures. Many MS instruments and data acquisition strategies are available, and the data they acquire differ substantially, thus requiring tailored analysis algorithms. Hence, many dedicated bioinformatics workflows…
Yi Lyu, Pei-Chieh Lo, Natan Lidukhover
The problem that we are trying to solve is to build a new benchmark for the data lakes. Although there is an increasing need for data lakes these days, there are not many standardized benchmarks for data lakes. By doing that, we are hoping to get a more objective and comparative understanding of different data lake…
Andrew M. Bean, Kearns, Ryan Othniel, Angelika Romanou + 41 more
Jan Batzner3, 4 Negar Foroutan 2 Chris Schmitz 5 Karolina Korgul 1 Hunar Batra 1 Oishi Deb 1 Emma Beharry 6 Cornelius Emde 1 Thomas Foster 1 Anna Gausen 7 María Grandury8, 9 Simeng Han 10 Valentin Hofmann11, 12 Lujain Ibrahim 1 Hazel Kim 1 Hannah Rose Kirk1, 7 Fangru Lin 1 Gabrielle Kaili-May Liu 10 Lennart Luettgau 7…
Izaskun Mallona, Charlotte Soneson, Ben Carrillo, Almut Lütge + 5 more
Benchmarking, which involves collecting reference datasets and demonstrating method performance, is a requirement for the development of new computational tools, but also becomes a domain of its own to achieve neutral comparisons of methods. Although a lot has been written about how to design and conduct benchmark…
Zhihong Zhan, Maolin Ye, Michael C. Orr, Weiqiang Chen + 4 more
Biodiversity can be quantified only after organisms are assigned to reproducible units, yet most individuals encountered in nature lack reliable species-level identifications. Molecular operational taxonomic units can organize unnamed diversity, whereas image-based approaches generally depend on predefined species…
Authors not listed
Accurately predicting chemical reaction yields in silico is a long-standing goal in organic chemistry that, if achieved, would revolutionize synthesis design, op-timization, and discovery. The vast reaction data within scientific literature rep-resents a rich resource for training predictive machine learning models…
Authors not listed
We present the next generation of AMP, a neural network potential (NNP) with anisotropic message passing designed to study large biomolecular systems at DFT accuracy in the condensed phase using a multiscale approach similar to quantum-mechanics/molecular-mechanics (QM/MM) with electrostatic embedding. We trained AMPv3…
Benton Chuter, Min Young Kim, Andrew B. Stiemke, Nikhil Dave + 6 more
To systematically review automated nerve morphometry tools and independently benchmark their performance on independent optic nerve datasets. Systematic review and comparative benchmarking study. Benchmarking was performed using paraphenylenediamine-stained mouse (n = 85) and rat (n = 44) optic nerve images with…
Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman + 14 more
The rapid scaling of Genomic Foundation Models (GFMs) has created a critical need for standardized evaluation frameworks. Current benchmarking practices are often fragmented, relying on model-specific preprocessing and inconsistent metric implementations that hinder reproducible comparisons. We present GFMBench-API, a…
Isaac Corley, Nils Lehmann, Caleb Robinson, Gabriel Tseng + 5 more
Geospatial foundation models (GFMs) have been proposed as generalizable backbones for disaster response, land-cover mapping, food-security monitoring, and other high-stakes Earth-observation tasks. Yet the published work about these models does not give reviewers or users enough information to tell which model fits a…
Yiqun Chen, Stephanie C. Hicks
Scientific coding agents are difficult to benchmark because many research tasks require executable work yet produce ambiguous or hard-to-verify outputs. Because benchmark construction requires substantial time and resources, automation offers a path to accelerating methods evaluation. We introduce an interactive…