27 papers · ranked by Valyu relevance
Davide Chicco, Giuseppe Jurman
Background To evaluate binary classifications and their confusion matrices, scientific researchers can employ several statistical rates, accordingly to the goal of the experiment they are investigating. Despite being a crucial issue in machine learning, no widespread consensus has been reached on a unified elective…
Giles M. Foody, Shigao Huang
The accuracy of a classification is fundamental to its interpretation, use and ultimately decision making. Unfortunately, the apparent accuracy assessed can differ greatly from the true accuracy. Mis-estimation of classification accuracy metrics and associated mis-interpretations are often due to variations in…
Margherita Grandini, E. Bagli, Giorgio Visani
Classification tasks in machine learning involving more than two classes are known by the name of "multi-class classification". Performance indicators are very useful when the aim is to evaluate and compare different classification models or machine learning techniques. Many metrics come in handy to test the ability of…
Silvia Beddar-Wiesing, Alice Moallemy-Oureh, Marie Kempkes, Josephine M. Thomas
Machine Learning is a diverse field applied across various domains such as computer science, social sciences, medicine, chemistry, and finance. This diversity results in varied evaluation approaches, making it difficult to compare models effectively. Absolute evaluation measures offer a practical solution by assessing…
Francisco J. Valverde-Albacete, Carmen Peláez-Moreno, Matteo G. A. Paris
'Matteo G. A. Paris'] The most widely spread measure of performance, accuracy, suffers from a paradox: predictive models with a given level of accuracy may have greater predictive power than models with higher accuracy. Despite optimizing classification error rate, high accuracy models may fail to capture crucial…
Hooman H. Rashidi, Samer Albahra, Scott Robertson, Nam K. Tran + 1 more
'Bo Hu'] One of the core elements of Machine Learning (ML) is statistics and its embedded foundational rules and without its appropriate integration, ML as we know would not exist. Various aspects of ML platforms are based on statistical rules and most notably the end results of the ML model performance cannot be…
Areen Arabiat, Hamza Abu Owida, Suhaila Abuowaida, Nawaf Alshdaifat + 2 more
This study emphasizes the potential of computational techniques in cancer risk assessment, highlighting opportunities for specific and data-driven healthcare solutions. It examines the use of artificial intelligence (AI), machine learning (ML), and deep learning (DL) approaches to improve cancer risk assessment using a…
Richard Dinga, Brenda W.J.H. Penninx, Dick J. Veltman, Lianne Schmaal + 1 more
Pattern recognition predictive models have become an important tool for analysis of neuroimaging data and answering important questions from clinical and cognitive neuroscience. Regardless of the application, the most commonly used method to quantify model performance is to calculate prediction accuracy, i.e. the…
Yuli Slavutsky, Yuval Benjamini
Multiclass classifiers are often designed and evaluated only on a sample from the classes on which they will eventually be applied. Hence, their final accuracy remains unknown. In this work we study how a classifier's performance over the initial class sample can be used to extrapolate its expected accuracy on a…
Vincent Labatut, Hocine Cherifi
— The selection of the best classification algorithm for a given dataset is a very widespread problem. It is also a complex one, in the sense it requires to make several important methodological choices. Among them, in this work we focus on the measure used to assess the classification performance and rank the…
Jan Kozak, Barbara Probierz, Krzysztof Kania, Przemysław Juszczuk + 1 more
Classification is one of the main problems of machine learning, and assessing the quality of classification is one of the most topical tasks, all the more difficult as it depends on many factors. Many different measures have been proposed to assess the quality of the classification, often depending on the application…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European Chemicals Agency (ECHA) and US Environmental Protection Agency (EPA) have listed approximately 800k chemicals that must be further investigated for their potential environmental and/or human health risk. A significant number of these chemicals have large enough global volumes of consumption (e.g.…
Mario Franco, Gerardo L. Febres, Nelson Fernández, Carlos Gershenson
Classification is a ubiquitous and fundamental problem in artificial intelligence and machine learning, with extensive efforts dedicated to developing more powerful classifiers and larger datasets. However, the classification task is ultimately constrained by the intrinsic properties of datasets, independently of…
O.C. Metcalf, C. Alencar Nunes, W.A. Hopping, A.C. Lees + 2 more
Automated detection and classification of species vocalisations offers the potential to utilise acoustic datasets across unprecedented spatial and temporal scales. However, classification algorithms inevitably generate errors, and error rates vary with context. While methods for quantifying error rates in ecoacoustics…
Sung-Cheol Kim, Adith S. Arun, Mehmet Eren Ahsen, Robert Vogel + 1 more
'Gustavo Stolovitzky'] Title: Significance While it would be desirable that the output of binary classification algorithms be the probability that the classification is correct, most algorithms do not provide a method to calculate such a probability. We propose a probabilistic output for binary classifiers based on an…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…
Seyyed Mahmood Ghasem, Johannes F. Fahrmann, Samir Hanash, Kim-Anh Do + 2 more
Logistic regression has demonstrated its utility in classifying binary labeled datasets through the maximum likelihood approach. However, in numerous biological and clinical contexts, the aim is often to determine coefficients that yield the highest sensitivity at the pre-specified specificity or vice versa. Therefore…
James Wellnitz, Sankalp Jain, Joshua Hochuli, Travis Maxfield + 3 more
Traditional best practices for Quantitative Structure Activity Relationship (QSAR) modeling recommend dataset balancing and balanced accuracy (BA) as the key desired objective of model development. This study challenges the conventional norms by recommending the use of models with the highest positive predictive value…
Areej Fatemah Meghji, Naeem Ahmed Mahoto, Yousef Asiri, Hani Alshahrani + 3 more
'Hani Alshahrani' 'Adel Sulaiman' 'Asadullah Shaikh' 'Shadi Aljawarneh'] Higher educational institutes generate massive amounts of student data. This data needs to be explored in depth to better understand various facets of student learning behavior. The educational data mining approach has given provisions to extract…
Telmo M. Silva Filho, Hao Song, Miquel Perelló-Nieto, Raúl Santos‐Rodríguez + 2 more
'Raúl Santos‐Rodríguez' 'Meelis Kull' 'Peter Flach'] This paper provides both an introduction to and a detailed overview of the principles and practice of classifier calibration. A well-calibrated classifier correctly quantifies the level of uncertainty or confidence associated with its instance-wise predictions. This…
Niklas Tötsch, Daniel Hoffmann
Classifiers are often tested on relatively small data sets, which should lead to uncertain performance metrics. Nevertheless, these metrics are usually taken at face value. We present an approach to quantify the uncertainty of classification performance metrics, based on a probability model of the confusion matrix.…
Jonathan Fine, Anand Rasjashekar, Krupal P. Jethava, Gaurav Chopra
State-of-the-art identification of the functional groups present in an unknown chemical entity requires expertise of a skilled spectroscopist to analyse and interpret Fourier Transform Infra-Red (FTIR), Mass Spectroscopy (MS) and/or Nuclear Magnetic Resonance (NMR) data. This process can be time-consuming and…
Preston Raab, W. Evan Johnson, Stephen R. Piccolo
Precision medicine relies on accurate and generalizable predictions for patients across the spectrum of human diversity. Because capturing biological heterogeneity requires large sample sizes, researchers must often aggregate data from several experimental batches or independent studies. This integration allows for…
Muhammad Hanzla, Abdul Rehman Shinwari
Machine Learning (ML) can be defined as a class of Artificial Intelligence for automated data analysis, which is capable of detecting patterns in data. The extracted patterns can be used to predict un-known data or to assist in decision-making processes under uncertainty. Recent advances in experimental and…
Authors not listed
This study presents a validation and refinement of the “yellow cards” error detection workflow that can be applied to any property connected to molecular structure. In our implementation the workflow employed 5 predictive models with each assigning a “yellow card” to 5% of the entries with worst prediction accuracy.…
Jonathan A Fine, Judy Kuan-Yu Liu, Armen Beck, Kawthar Alzarieni + 4 more
Diagnostic ion-molecule reactions using tandem mass spectrometry can differentiate between isomeric compounds unlike a popular collision-activated dissociation methodology for the identification of previously unknown mixtures. Selected neutral reagents, such as 2-methoxypropene (MOP) are introduced into an ion trap…