24 papers · ranked by Valyu relevance
Nives Škunca, Adrian Altenhoff, Christophe Dessimoz, Lars Juhl Jensen
'Lars Juhl Jensen'] Gene Ontology (GO) has established itself as the undisputed standard for protein function annotation. Most annotations are inferred electronically, i.e. without individual curator supervision, but they are widely considered unreliable. At the same time, we crucially depend on those automated…
Svetlana Kiritchenko, Saif M. Mohammad
Rating scales are a widely used method for data annotation; however, they present several challenges, such as difficulty in maintaining inter- and intra-annotator consistency. Best–worst scaling (BWS) is an alternative method of annotation that is claimed to produce high-quality annotations while keeping the required…
Owen Cook, Jake Vasilakes, Ian Roberts, Xingyi Song
Data annotation is an essential component of the machine learning pipeline; it is also a costly and time-consuming process. With the introduction of transformer-based models, annotation at the document level is increasingly popular; however, there is no standard framework for structuring such tasks. The EffiARA…
Carlos A. Martínez-Miwa, Mario Castelán
We documented the relabeling process for a subset of a renowned database for emotion-in-context recognition, with the aim of promoting reliability in final labels. To this end, emotion categories were organized into eight groups, while a large number of participants was requested for tagging. A strict control strategy…
Reid Swanson, Stephanie M. Lukin, Luke Eisenberg, Thomas Corcoran + 1 more
'Marilyn Walker'] The language used in online forums differs in many ways from that of traditional language resources such as news. One difference is the use and frequency of nonliteral, subjective dialogue acts such as sarcasm. Whether the aim is to develop a theory of sarcasm in dialogue, or engineer automatic…
Daniel J. Miller, Blaise Gratton, Zoe LeBlanc, Jon H. Kaas
Supervised statistical learning for cell-level segmentation and morphometry in optical microscopy is limited less by algorithmic capacity than by the scarcity of reliable, expert-validated ground truth. In comparative neuroscience and quantitative histology, where classical stains such as Nissl’s method remain the…
Oliver Cook, Charlie Grimshaw, Ben Wu, Sophie Dillon + 5 more
Knowledge-Based Misinformation Detection on Social Media Authors: ['Oliver Cook' 'Charlie Grimshaw' 'Ben Wu' 'Sophie Dillon' 'Jack M. Hicks' 'Luke Jones' 'T Smith' 'Matyas Szert' 'Xingyi Song'] Misinformation spreads rapidly on social media, confusing the truth and targetting potentially vulnerable people. To…
Kristina Yordanova, Frank Krüger
Providing ground truth is essential for activity recognition and behaviour analysis as it is needed for providing training data in methods of supervised learning, for providing context information for knowledge-based methods, and for quantifying the recognition performance. Semantic annotation extends simple symbolic…
Brandon M. Booth, Shrikanth S. Narayanan
Accurately representing changes in mental states over time is crucial for understanding their complex dynamics. However, there is little methodological research on the validity and reliability of human-produced continuous-time annotation of these states. We present a psychometric perspective on valid and reliable…
Federico Cabitza, Andrea Campagner, Luca Maria Sconfienza
Background We focus on the importance of interpreting the quality of the labeling used as the input of predictive models to understand the reliability of their output in support of human decision-making, especially in critical domains, such as medicine. Methods Accordingly, we propose a framework distinguishing the…
Jan-Christoph Klie, Richard Eckart de Castilho, Iryna Gurevych
Data quality is crucial for training accurate, unbiased, and trustworthy machine learning models as well as for their correct evaluation. Recent works, however, have shown that even popular datasets used to train and evaluate state-of-the-art models contain a non-negligible amount of erroneous annotations, biases, or…
Melanie J Martin
Background In this paper we present a detailed scheme for annotating medical web pages designed for health care consumers. The annotation is along two axes: first, by reliability (the extent to which the medical information on the page can be trusted), second, by the type of page (patient leaflet, commercial, link…
Oana Inel, Tim Draws, Lora Aroyo
The rapid entry of machine learning approaches in our daily activities and high-stakes domains demands transparency and scrutiny of their fairness and reliability. To help gauge machine learning models' robustness, research typically focuses on the massive datasets used for their deployment, e.g., creating and…
Authors not listed
As the volume and diversity of bioactivity data in ChEMBL continues to grow, ensuring that assay metadata is standardized, interoperable, and machine-readable is critical for effective use in cheminformatics and ML applications. In this work, we present recent efforts to enhance the quality and granularity of bioassay…
Authors not listed
Mass spectrometry-based natural products targeted discovery often relies on a complicated decision-making process involving tedious comparison of exact masses data and tandem mass spectra-based annotation tools output against various spectral reference libraries. To address this bottleneck, we present tandem mass…
Authors not listed
Computational models predicting the sites of metabolism (SOM) of small or- ganic molecules have become invaluable tools for studying and optimizing the metabolic properties of xenobiotics. However, the performance of SOM predic- tors has shown signs of plateauing in recent years, primarily due to the limited…
Granville J. Matheson
Positron emission tomography (PET), along with many other fields of clinical research, is both timeconsuming and expensive, and recruitable patients can be scarce. These constraints limit the possibility of large-sample experimental designs, and often lead to statistically underpowered studies. This problem is…
Wout Bittremieux, Mingxun Wang, Pieter C. Dorrestein
Background: Spectral library searching is currently the most common approach for compound annotation in untargeted metabolomics. Spectral libraries applicable to liquid chromatography mass spectrometry have grown in size over the past decade to include hundreds of thousands to millions of mass spectra and tens of…
Authors not listed
The exponential growth of chemical literature necessitates the development of automated tools for extracting and curating molecular information from unstructured scientific publications into open-access chemical databases. Current optical chemical structure recognition (OCSR) and named entity recognition solutions…
Keiji Ota, Anthony Ciston, Patrick Haggard, Thibault Gajdos Preuss + 1 more
A key challenge in today’s fast-paced digital world is to integrate information from various sources, which differ in their reliability. Yet, little is known about how explicit probabilistic information on the likelihood of a source to provide correct information is used in decision-making. Here, we investigated how…
Laurie Compère, Greg J. Siegle, Kymberly Young
Proponents of personalized medicine have promoted neuroimaging evaluation and treatment of major depressive disorder in three areas of clinical application: clinical prediction, outcome evaluation, and neurofeedback. Whereas psychometric considerations such as test-retest reliability are basic precursors to clinical…
Keiji Ota, Anthony Ciston, Patrick Haggard, Thibault Gajdos Preuss + 1 more
A key challenge in today’s fast-paced digital world is to integrate information from various sources, which differ in their reliability. Yet, little is known about how explicit probabilistic information about the likelihood that a source provides correct information is used in decision-making. Here, we investigated how…
Friedrich Hastedt, Rowan M. Bailey, Klaus Hellgardt, Sophia N. Yaliraki + 2 more
Machine learning models for chemical retrosynthesis have attracted substantial interest in recent years. Unaddressed challenges, particularly the absence of robust evaluation metrics for performance comparison, and the lack of black-box interpretability, obscure model limitations and impede progress in the field. We…
Authors not listed
Machine learning (ML) models are increasingly used in quantum chemistry, but their reliability hinges on uncertainty quantification (UQ). In this study, we compare two prominent UQ paradigms—Deep Evidential Regression (DER) and Deep Ensembles—on the QM9 and WS22 datasets, with a specific emphasis on the role of post…