22 papers · ranked by Valyu relevance
Yiling Chen, Shi Feng, Paul Kattuman, Fang-Yi Yu
How can we assess the reliability of a dataset without access to ground truth? We introduce the problem of reliability scoring for datasets collected from potentially strategic sources. The true data are unobserved, but we see outcomes of an unknown statistical experiment that depends on them. To benchmark reliability…
Sharmila Vaz, Torbjörn Falkmer, Anne Elizabeth Passmore, Richard Parsons + 2 more
'Richard Parsons' 'Pantelis Andreou' 'Susanne Hempel'] The use of standardised tools is an essential component of evidence-based practice. Reliance on standardised tools places demands on clinicians to understand their properties, strengths, and weaknesses, in order to interpret results and make clinical decisions.…
Xiuxiu Tang, G. Alex Ambrose, Ying Cheng
Student responses in STEM assessments are often handwritten and combine symbolic expressions, calculations, and diagrams, creating substantial variation in format and interpretation. Despite their importance for evaluating students' reasoning, such responses are time-consuming to score and prone to rater inconsistency…
C. Berthomier, V. Muto, C. Schmidt, G. Vandewalle + 13 more
New challenges in sleep science require to describe fine grain phenomena or to deal with large datasets. Beside the human resource challenge of scoring huge datasets, the inter- and intra-expert variability may also reduce the sensitivity of such studies. Searching for a way to disentangle the variability induced by…
Granville J. Matheson, Andrew Gray
Neuroimaging, in addition to many other fields of clinical research, is both time-consuming and expensive, and recruitable patients can be scarce. These constraints limit the possibility of large-sample experimental designs, and often lead to statistically underpowered studies. This problem is exacerbated by the use of…
Granville J. Matheson
Positron emission tomography (PET), along with many other fields of clinical research, is both timeconsuming and expensive, and recruitable patients can be scarce. These constraints limit the possibility of large-sample experimental designs, and often lead to statistically underpowered studies. This problem is…
Robert M. Talbot
Score reliability is necessary for establishing a validity argument for an instrument, and is therefore highly important to investigate. Depending on the proposed instrument use and score interpretations, differing degrees of precision in measurement or reliability are required. Researchers sometimes fail to take a…
Étienne Marcotte, Valentina Zantedeschi, Alexandre Drouin, Nicolas Chapados
'Nicolas Chapados'] Multivariate probabilistic time series forecasts are commonly evaluated via proper scoring rules, i.e., functions that are minimal in expectation for the ground-truth distribution. However, this property is not sufficient to guarantee good discrimination in the non-asymptotic regime. In this paper…
Alexander Muacevic, John R Adler, Veena K Ranganath, Ami Ben-Artzi + 8 more
'Jenny Brook' 'Yosra Suliman' 'Astrid Floegel-Shetty' 'Thasia Woodworth' 'Mihaela Taylor' 'Laurie A Ramrattan' 'David Elashoff' 'Gurjit S Kaeley'] Objective: Musculoskeletal ultrasound real-time image acquisition and scoring are complex, and many factors affect reliability. Static image reliability does not guarantee…
Peter M. Visscher, Loic Yengo
In this study we quantify the accuracy of scoring the quality of research grants using a finite set of distinct categories (1, 2, …., k), when the unobserved grant score is a continuous random variable comprising a true quality score and measurement error, both normally distributed. We vary the number of categories…
Dario Cecilio-Fernandes, Harro Medema, Carlos Fernando Collares, Lambert Schuwirth + 2 more
'Lambert Schuwirth' 'Janke Cohen-Schotanus' 'René A. Tio'] Background Progress testing is an assessment tool used to periodically assess all students at the end-of-curriculum level. Because students cannot know everything, it is important that they recognize their lack of knowledge. For that reason, the formula-scoring…
Tatsuya Yoshizawa, Shoichi Ishida, Tomohiro Sato, Masateru Ohta + 2 more
Molecular design using data-driven generative models has emerged as a promising technology, impacting various fields such as drug discovery and the development of functional materials. However, this approach is often susceptible to optimization failure due to reward hacking, where prediction models fail to accurately…
Bruno Kopp, Florian Lange, Alexander Steinke
The Wisconsin Card Sorting Test (WCST) represents the gold standard for the neuropsychological assessment of executive function. However, very little is known about its reliability. In the current study, 146 neurological inpatients received the Modified WCST (M-WCST). Four basic measures (number of correct sorts…
Martin Gell, Mauricio S. Hoffmann, Tyler M. Moore, Aki Nikolaidis + 8 more
Identifying robust brain-psychopathology associations with neuroimaging remains difficult, in part due to substantial heterogeneity within and comorbidity between diagnostic categories. Transdiagnostic latent factor models aim to address this structure by separating shared and unique symptom variance. However, it…
Jan Kadlec, Catherine Walsh, Uri Sadé, Ariel Amir + 2 more
The surge in interest in individual differences has coincided with the latest replication crisis centered around brain-wide association studies of brain-behavior correlations. Yet the reliability of the measures we use in cognitive neuroscience, a crucial component of this brain-behavior relationship, is often assumed…
Gang Chen, Daniel S. Pine, Melissa A. Brotman, Ashley R. Smith + 2 more
The concept of test-retest reliability indexes the consistency of a measurement across time. High reliability is critical for any scientific study, but specifically for the study of individual differences. Evidence of poor reliability of commonly used behavioral and functional neuroimaging tasks is mounting. Reports on…
Authors not listed
Graph Neural Networks (GNNs) are powerful tools for molecular property prediction, but they are not magic. When applied to molecules unlike their training data, they produce unreliable predictions that are difficult to detect. The Applicability Domain (AD) concept addresses this by defining regions of chemical space…
Authors not listed
Ensuring the trustworthiness of machine learning (ML) models in high-stake applications is crucial. One such application is predicting anti-cancer drug sensitivity, where ML models are built with the final goal of integrating them into treatment recommendation systems for personalized medicine. Here, we propose a…
Laure Ciernik, Agnieszka Kraft, Florian Barkmann, Josephine Yates + 1 more
In the field of single-cell RNA sequencing (scRNA-seq), gene signature scoring is integral for pinpointing and characterizing distinct cell populations. However, challenges arise in ensuring the robustness and comparability of scores across various gene signatures and across different batches and conditions. Here, we…
Hailiang Du
The evaluation of probabilistic forecasts plays a central role both in the interpretation and in the use of forecast systems and their development. Probabilistic scores (scoring rules) provide statistical measures to assess the quality of probabilistic forecasts. Often, many probabilistic forecast systems are available…
Fergus Boyles, Charlotte M Deane, Garrett Morris
Machine learning scoring functions for protein-ligand binding affinity prediction have been found to consistently outperform classical scoring functions. Structure-based scoring functions for universal affinity prediction typically use features describing interactions derived from the protein-ligand complex, with…
Yasmine Nahal, Janosch Menke, Julien Martinelli, Markus Heinonen + 5 more
Machine learning (ML) systems have enabled the modelling of quantitative structure-property relationships (QSPR) and structure-activity relationships (QSAR) using existing experimental data to predict target properties for new molecules. These property predictors hold significant potential in accelerating drug…