25 papers · ranked by Valyu relevance
Lingyan Xue, Xinyu Zhang, Weidong Jiang, Kai Huo
Machine learning and deep learning classification models are data-driven, and the model and the data jointly determine their classification performance. It is biased to evaluate the model's performance only based on the classifier accuracy while ignoring the data separability. Sometimes, the model exhibits excellent…
Jan Kozak, Barbara Probierz, Krzysztof Kania, Przemysław Juszczuk + 1 more
Classification is one of the main problems of machine learning, and assessing the quality of classification is one of the most topical tasks, all the more difficult as it depends on many factors. Many different measures have been proposed to assess the quality of the classification, often depending on the application…
Akshay Akshay, Masoud Abedi, Navid Shekarchizadeh, Fiona C. Burkhard + 5 more
A performance metric is a tool to measure the correctness of a trained Machine Learning (ML) model. Numerous performance metrics have been developed for classification problems making it overwhelming to select the appropriate one since each of them represents a particular aspect of the model. Furthermore, selection of…
David J. Hand, Peter Christen, Sumayya Ziyad
the problem Authors: ['David J. Hand' 'Peter Christen' 'Sumayya Ziyad'] The problem of identifying to which of a given set of classes objects belong is ubiquitous, occurring in many research domains and application areas, including medical diagnosis, financial decision making, online commerce, and national security.…
Ningsheng Zhao, Trang Bui, Jia Yuan Yu, Krzysztof Dzieciolowski
Many classification performance metrics exist, each suited to a specific application. However, these metrics often differ in scale and can exhibit varying sensitivity to class imbalance rates in the test set. As a result, it is difficult to use the nominal values of these metrics to interpret and evaluate…
Yeliz Senkaya, Cetin Kurnaz, Ferdi Ozbilgin, Yong-An Chung
Background/Objectives: Alzheimer’s disease (AD) is a devastating neurodegenerative disorder that progressively impairs cognitive, neurological, and behavioral functions, severely affecting quality of life. The current diagnostic process relies on expert interpretation of extensive clinical assessments, often leading to…
Nureni Ayofe Azeez, Sanjay Misra, Davidson Onyinye Ogaraku, Ademola Philip Abidoye + 2 more
The pervasive spread of fake news in online social media has emerged as a critical threat to societal integrity and democratic processes. To address this pressing issue, this research harnesses the power of supervised AI algorithms aimed at classifying fake news with selected algorithms. Algorithms such as Passive…
Siddharth Chaini, A. Mahabal, Ajit Kembhavi, Federica Bianco
aDepartment of Physics and Astronomy, University of Delaware, Newark, DE 19716, USA bDepartment of Physics, Indian Institute of Science Education and Research, Bhopal 462066, India cDivision of Physics, Mathematics and Astronomy, California Institute of Technology, Pasadena, CA 91125, USA dCenter for Data Driven…
Areen Arabiat, Hamza Abu Owida, Suhaila Abuowaida, Nawaf Alshdaifat + 2 more
This study emphasizes the potential of computational techniques in cancer risk assessment, highlighting opportunities for specific and data-driven healthcare solutions. It examines the use of artificial intelligence (AI), machine learning (ML), and deep learning (DL) approaches to improve cancer risk assessment using a…
Jay Devine, Helen K. Kurki, Jonathan R. Epp, Paula N. Gonzalez + 2 more
Classification is a fundamental task in biology used to assign members to a class. While linear discriminant functions have long been effective, advances in phenotypic data collection are yielding increasingly high-dimensional datasets with more classes, unequal class covariances, and non-linear distributions. Numerous…
Antonio García-Domínguez, Carlos E. Galván-Tejada, Rafael Magallanes-Quintanar, Hamurabi Gamboa-Rosales + 3 more
'Rafael Magallanes-Quintanar' 'Hamurabi Gamboa-Rosales' 'Irma González Curiel' 'Jesús Peralta-Romero' 'Miguel Cruz'] The development of medical diagnostic models to support healthcare professionals has witnessed remarkable growth in recent years. Among the prevalent health conditions affecting the global population…
Hajo Holzmann, Bernhard Klar
We show that established performance metrics in binary classification, such as the F-score, the Jaccard similarity coefficient or Matthews' correlation coefficient (MCC), are not robust to class imbalance in the sense that if the proportion of the minority class tends to 0, the true positive rate (TPR) of the Bayes…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Yuting Yang, Golrokh Mirzaei, Muhammad Umer
Cancer, in any of its forms, remains a significant public health concern worldwide. Advances in early detection and treatment could lead to a decline in the overall death rate from cancer in recent decades. Therefore, tumor prediction and classification play an important role in fighting cancer. This study built…
Philipp Thölke, Yorguin Jose Mantilla Ramos, Hamza Abdelhedi, Charlotte Maschke + 10 more
Machine learning (ML) is becoming a standard tool in neuroscience and neuroimaging research. Yet, because it is such a powerful tool, the appropriate application of ML requires a sound understanding of its subtleties and limitations. In particular, applying ML to datasets with imbalanced classes, which are very common…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
Asif Newaz, Artur B Adib, Taskeed Jabid
Framework for Imbalanced Classification Authors: ['Asif Newaz' 'Artur B Adib' 'Taskeed Jabid'] - Cost-sensitive learning is a popular technique used in the imbalanced domain. - Higher misclassification costs are naively applied to all minority-class instances. - In the proposed approach, instances are penalized…
Areej Fatemah Meghji, Naeem Ahmed Mahoto, Yousef Asiri, Hani Alshahrani + 3 more
'Hani Alshahrani' 'Adel Sulaiman' 'Asadullah Shaikh' 'Shadi Aljawarneh'] Higher educational institutes generate massive amounts of student data. This data needs to be explored in depth to better understand various facets of student learning behavior. The educational data mining approach has given provisions to extract…
Henri Tiittanen, Liisa Holm, Petri Törönen
Automated protein Function Prediction (AFP) is an intensively studied topic. Most of this research focuses on methods that combine multiple data sources, while fewer articles look for the most efficient ways to use a single data source. Therefore, we wanted to test how different preprocessing methods and classifiers…
Prashanth Athri, Vidhya Murali, Pradyumna Y Muralidhar, Cassandra Königs + 4 more
- 1. Department of Computer Science and Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Bengaluru, India - 2. PES Center for Pattern Recognition, Department of Computer Science and Engineering, PES University, Bengaluru, India - 3. Bioinformatics and Medical Informatics, Bielefeld University…
Chi Zhang, Dmytro Antypov, Matthew J Rosseinsky, Matthew Stephen Dyer
Machine learning has found wide application in the materials field, particularly in discovering structure-property relationships. However, its potential in predicting synthetic accessibility of materials remains relatively unexplored due to the lack of negative data. In this study, we employ several one-class…
Authors not listed
Early-stage drug discovery often suffers from data scarcity and out-of-distribution (OOD) shifts, which constrain the reliability of predictive models. While deep learning has advanced representation learning from molecular and biological data, tabular modeling remains indispensable, particularly in small-sample and…
Esteban Bertsch Aguilar, Sebastián Suñer Sánchez, Silvana Pinheiro, William J. Zamora Ramírez
- 1. 1. CBio3 Laboratory, School of Chemistry, University of Costa Rica, San Pedro, San José, Costa Rica - 2. 2. Laboratory of Computational Toxicology and Artificial Intelligence (LaToxCIA), Biological Testing Laboratory (LEBi), University of Costa Rica, San Pedro, San José, Costa Rica - 3. 3. Advanced Computing Lab…
Mohammad Ali Nematollahi, Soodeh Jahangiri, Arefeh Asadollahi, Maryam Salimi + 9 more
'Maryam Salimi' 'Azizallah Dehghan' 'Mina Mashayekh' 'Mohamad Roshanzamir' 'Ghazal Gholamabbas' 'Roohallah Alizadehsani' 'Mehdi Bazrafshan' 'Hanieh Bazrafshan' 'Hamed Bazrafshan drissi' 'Sheikh Mohammed Shariful Islam'] We used machine learning methods to investigate if body composition indices predict hypertension.…