19 papers · ranked by Valyu relevance
Angela Jin, Alexander Asemota, Dan E. Krane, Nathaniel D. Adams + 1 more
AI governance efforts increasingly rely on audit standards: agreed-upon practices for conducting audits. However, poorly designed standards can hide and lend credibility to inadequate systems. We explore how an audit standard's design influences its effectiveness through a case study of ASB 018, a standard for auditing…
Aline Gonçalves Capella, Marta Martín López, Juan José de Damborenea González, María Ángeles Arenas + 1 more
Laser-induced breakdown spectroscopy (LIBS) is a versatile technique for characterizing materials and analyzing elements, but it is limited in its quantitative accuracy due to nonlinear effects. This study uses machine learning (ML) regression algorithms to improve the quantification of commercially pure aluminum and…
Mohsen Sadatsafavi, Paul Gustafson, Solmaz Setayeshgar, Laure Wynants + 1 more
Contemporary sample size calculations for external validation of risk prediction models require users to specify fixed values of assumed model performance metrics alongside target precision levels (e.g., 95% CI widths). However, due to the finite samples of previous studies, our knowledge of true model performance in…
Freiesleben, Timo, Zezulka, Sebastian
Predictive benchmarking, the evaluation of machine learning models based on predictive performance and competitive ranking, is a central epistemic practice in machine learning research and an increasingly prominent method for scientific inquiry. Yet, benchmark scores alone provide at best measurements of model…
Authors not listed
Transition-state (TS) identification for bimolecular liquid-phase reactions is notoriously sensitive to the initial spatial arrangement of reactants, making automated searches difficult, especially in solvation where conformational effects dominate barrier heights. We address this gap with a fully automated, heuristic…
Alastair Pickering, Santiago Martinez Balvanera, Nicholas Brown, Sareach Chea + 4 more
1. Passive acoustic monitoring (PAM) is increasingly used for ecological research, biodiversity monitoring, assessment, and reporting. Automated species classifiers make it feasible to process large audio datasets but generate numerous detections that often need validation before use in downstream analyses or formal…
Authors not listed
We present graphRC, a graph-based method for rapid transition state (TS) mode analysis that provides chemical insight along normal mode displacements and reaction coordinate trajectories by translating Cartesian displacements into meaningful internal coordinate changes. Internal coordinates are constructed using…
Lara L. Russell-Lasalandra, Alexander P. Christensen, Hudson Golino
To demonstrate the validity of the surveys produced by AI-GENIE, we created five new Big Five personality surveys using Gemma 2, GPT-3.5, GPT-4o, Llama 3, and Mixtral. The same prompt and instructions were used for each model (the full prompt is provided in the [Sec39]). Each model generated at least 40 items using a…
Michael Biddle, Jemma Cooper, Katherine Blades, Dominic Ruddy + 2 more
Lack of antibody validation by researchers frequently misdirects biomedical research, yet the ethical consequences — particularly avoidable use of animal and human biological materials — remain unquantified. Using focus groups (n=12), surveys (n=107), and systematic analysis of 785 publications, we examined how…
Kristof Hofrichter, Lukas Elster, Clemens Linnhoff, Timm Ruppert + 2 more
Simulation-based testing is playing an increasingly important role in the development and validation of automated driving functions, as real-world testing is often limited by cost, safety, and scalability. An essential part of this is the simulation of active perception sensors such as lidar and radar which enable…
Vincenzo Caretti, Eleonora Topino, Andrea Fontana, Gianluigi Di Cesare + 4 more
The internal saboteur may be understood as a multidimensional configuration of maladaptive inner processes involving recurrent negative self-evaluation, distressing relational expectations, repetitive negative thinking, and self-undermining inner experiences. Within this framework, the present study aimed to develop…
Luis Miguel Rojo-Bofill, Juan Pablo Carrasco-Picazo, Amelia Rosa Granda-Pinan, Jose Martinez-Raga + 1 more
Background: Practical clinical training is a crucial part of undergraduate medical education. Assessing students’ satisfaction with this training is essential for improving education programmes. While research has often focused on student satisfaction with general or theoretical education, studies on practical clinical…
Romain Lefeuvre, Maïwenn Le Goasteller, Jessie Galasso, Benoit Combemale + 2 more
Empirical software engineering research often depends on datasets of code repository artifacts, where sampling strategies are employed to enable large-scale analyses. The design and evaluation of these strategies are critical, as they directly influence the generalizability of research findings. However, sampling…
Vishal Bharti, Debojyoti Chakraborty
Reproducible computational biology depends on statistical decisions that routine workflows often skip: verifying that a differential-expression test’s assumptions hold across all genes, that a strategy-comparison ANOVA is robust to non-normality, or that a meta-analysis is not distorted by publication bias. Surveys…
Orsolya Székely, Nicholas P. Holmes, Jennifer Ashton, Friederike Breuer + 13 more
We introduce the TMS-RAT, a reporting (assessment) tool for TMS studies Developed within a community-informed, iterative process rating 333 TMS studies Empirically evaluated for usability, inter-rater, and test-retest reliability A validated subset enables reliable retrospective assessment of reporting The modular…
Joao Pinelo, Joao Goncalves, Arun Shukla, Adriana Santos-Ferreira
The Internal Waves Service screens the Sentinel-1 Wave-mode archive for internal solitary waves, routing detections to experts whose adjudication time is the resource the effort exists to conserve. Because attention is the cost of error, precision leads. Its classifier was trained and reported at a one-to-one class…
Shuai Wang, Xinyuan Tian, Pangpang Liu, Yize Zhao
This paper argues that workflow closure is not scientific closure in auto-research systems. Current systems can increasingly complete research-like loops internally, moving from idea generation to experiment execution, writing, and self-evaluation. That achievement is real, but it does not by itself give the resulting…
R. Thacker
AlphaFold2 (AF2) has transformed structural biology, yet its confidence metrics, particularly the predicted Local Distance Difference Test (pLDDT) and predicted Template Modelling score (pTM), systematically fail for fold-switching proteins, which adopt two or more distinct conformations from a single amino acid…
Authors not listed
Transition state (TS) geometries of chemical reactions are key to understanding reaction mechanisms and estimating kinetic properties. Inferring these directly from 2D reaction graphs offers chemists a powerful tool for rapid and accessible reaction analysis. Quantum chemical methods for computing TSs are…