25 papers · ranked by Valyu relevance
Mark van der Loo, Edwin de Jonge
Data validation is the activity where one decides whether or not a particular data set is fit for a given purpose. Formalizing the requirements that drive this decision process allows for unambiguous communication of the requirements, automation of the decision process, and opens up ways to maintain and investigate the…
David A. Cook, Rose Hatala
Background Simulation plays a vital role in health professions assessment. This review provides a primer on assessment validation for educators and education researchers. We focus on simulation-based assessment of health professionals, but the principles apply broadly to other assessment approaches and topics. Key…
Gerard J. Kleywegt
The need for validation of macromolecular crystal structures is discussed. A general approach to validation is presented, together with examples of its implementation in the special case of macromolecular crystallography.
David Higgins, Christian Johner
The introduction of artificial intelligence / machine learning (AI/ML) products to the regulated fields of pharmaceutical research and development (R&D) and drug manufacture, and medical devices (MD) and in-vitro diagnostics (IVD), poses new regulatory problems: a lack of a common terminology and understanding leads to…
Sibel Eker, Elena Rovenskaya, Michael Obersteiner, Simon Langan
Quantitative modelling is commonly used to assist the policy dimension of sustainability problems. Validation is an important step to make models credible and useful. To investigate existing validation viewpoints and approaches, we analyse a broad academic literature and conduct a survey among practitioners. We find…
Ksenija Dvurecenska, Steve Graham, Edoardo Patelli, Eann A. Patterson
'Eann A. Patterson'] A new validation metric is proposed that combines the use of a threshold based on the uncertainty in the measurement data with a normalized relative error, and that is robust in the presence of large variations in the data. The outcome from the metric is the probability that a model's predictions…
Mark van der Loo, Edwin de Jonge
Checking data quality against domain knowledge is a common activity that pervades statistical analysis from raw data to output. The R package validate facilitates this task by capturing and applying expert knowledge in the form of validation rules: logical restrictions on variables, records, or data sets that should be…
K. Larsen, R. Lukyanenko, Roland M. Mueller, V. Storey + 3 more
Researchers must ensure that the claims about the knowledge produced by their work are valid. However, validity is neither well-understood nor consistently established in design science, which involves the development and evaluation of artifacts (models, methods, instantiations, and theories) to solve problems. As a…
Eric W. Deutsch, Roger Kramer, Joseph Ames, Andrew Bauman + 21 more
Translational biomedical research is generating exponentially more data: thousands of whole-genome sequences (WGS) are now available; brain data are doubling every two years. Analyses of Big Data, including imaging, genomic, phenotypic, and clinical data, present qualitatively new challenges as well as opportunities.…
Lucy Ellen Lwakatare, Ellinor Rånge, Ivica Crnković, Jan Bosch
—Background: Data errors are a common challenge in machine learning (ML) projects and generally cause significant performance degradation in ML-enabled software systems. To ensure early detection of erroneous data and avoid training ML models using bad data, research and industrial practice suggest incorporating a data…
Evgueni Jacob, Angélique Perrillat-Mercerot, Jean-Louis Palgen, Adèle L’Hostis + 5 more
Over the past several decades, metrics have been defined to assess the quality of various types of models and to compare their performance depending on their capacity to explain the variance found in real-life data. However, available validation methods are mostly designed for statistical regressions rather than for…
Zhe Wang, Ardan Patwardhan, Gerard J. Kleywegt
The Electron Microscopy Data Bank (EMDB) is the central archive of the electron cryo-microscopy (cryo-EM) community for storing and disseminating volume maps and tomograms. With input from the community, EMDB has developed new resources for validation of cryo-EM structures, focussing on the quality of the volume data…
Udit Surya Saha, Michele Vendruscolo, Anne E. Carpenter, Shantanu Singh + 2 more
Recent advances in machine learning methods for materials science have significantly enhanced accurate predictions of the properties of novel materials. Here, we explore whether these advances can be adapted to drug discovery by addressing the problem of prospective validation - the assessment of the performance of a…
Sterling Baird, Tran Diep, Taylor Sparks
We present Descending from Stochastic Clustering Variance Regression (DiSCoVeR), a Python tool for identifying high-performing, chemically unique compositions relative to existing compounds using a combination of a chemical distance metric, density-aware dimensionality reduction, and clustering. We introduce several…
Muaz A. Niazi, Amir Hussain, Mario Kolberg
—Agent Based Models are very popular in a number of different areas. For example, they have been used in a range of domains ranging from modeling of tumor growth, immune systems, molecules to models of social networks, crowds and computer and mobile self-organizing networks. One reason for their success is their…
Roy Eagleson, Leo Joskowicz, Nassir Navab, Philipp Fürnstahl + 3 more
'Mazda Farshad' 'Hooman Esfandiari' 'Matthias Seibold'] This paper presents a discussion about the fundamental principles of Analysis of Augmented and Virtual Reality (AR/VR) Systems for Medical Imaging and Computer-Assisted Interventions. The three key concepts of Analysis (Verification, Evaluation, and Validation)…
Olawale Salaudeen, Anka Reuel, Ahmed Ahmed, Suhana Bedi + 5 more
'Zachary Robertson' 'Sudharsan Sundar' 'Ben Domingue' 'Angelina Wang' 'Sanmi Koyejo'] While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning capabilities, are supported with model performance on…
Sterling Baird, Tran Diep, Taylor Sparks
We present Descending from Stochastic Clustering Variance Regression (DiSCoVeR), a Python tool for identifying high-performing, chemically unique compositions relative to existing compounds using a combination of a chemical distance metric, density-aware dimensionality reduction, and clustering. We introduce several…
Giuseppe Gallitto, Robert Englert, Balint Kincses, Raviteja Kotikalapudi + 4 more
Multivariate predictive models play a crucial role in enhancing our understanding of complex biological systems and in developing innovative, replicable tools for translational medical research. However, the complexity of machine learning methods and extensive data pre-processing and feature engineering pipelines can…
Victor H. R. Nogueira, Rishabh Sharma, Rafael V. C. Guido, Michael J. Keiser
As efforts to improve the robustness of molecular representations advance, so does the need for methods to test and validate them. We use a Variational Auto-Encoder (VAE), an unsupervised deep learning model, to generate anomalous samples of a well-known molecular string format called SELF-referencIng Embedded Strings…
Yanlin Zhang, Jing Zhao
We formalize scientific methodology—the end-to-end process from question formulation to evidence-grounded writing—as a phase-gated research protocol with explicit return paths and persistent constraints, and instantiate it for general-purpose language models as executable protocol specifications. The formalization…
Authors not listed
Approximately 40% of marketed drugs exhibit suboptimal pharmacokinetic profiles. Co-crystallization, where pairs of molecules form a multicomponent crystal, constitutes a promising strategy to enhance physicochemical properties without compromising the pharmacological activity. However, finding promising co-crystal…
Authors not listed
Automation of experiments in cloud laboratories promises to revolutionize scientific research by enabling remote experimentation and improving reproducibility. However, maintaining quality control without constant human oversight remains a critical challenge. Here, we present a novel machine learning framework for…
Authors not listed
This study presents a validation and refinement of the “yellow cards” error detection workflow that can be applied to any property connected to molecular structure. In our implementation the workflow employed 5 predictive models with each assigning a “yellow card” to 5% of the entries with worst prediction accuracy.…
Maria H. Rasmussen, Chenru Duan, Heather J. Kulik, Jan Halborg Jensen
With the increasingly more important role of machine learning (ML) models in chemical research, the need for putting a level of confidence to the model predictions naturally arises. Several methods for obtaining uncertainty estimates have been proposed in recent years but consensus on the evaluation of these have yet…