26 papers · ranked by Valyu relevance
Tiago Brogueira, Mário A. T. Figueiredo
Binary classification is one of the oldest, most prevalent and studied problems in machine learning. However, the metrics used to evaluate model performance have received comparatively little attention. The area under the receiver operating characteristic curve (AUROC) has long been a standard choice for model…
Ronaldo C. Prati
We propose a unified algebraic framework for classification performance evaluation that encompasses binary, multiclass, multilabel, ordinal, hierarchical, cost-sensitive, and soft-label settings within a single formalism. The foundation is a representation of actual and predicted labels as binary indicator matrices…
Dembélé, Doulaye
Classification is a machine learning method used in many practical applications: text mining, handwritten character recognition, face recognition, pattern classification, scene labeling, computer vision, natural langage processing. A classifier prediction results and training set information are often used to get a…
Xin Eric Wang, Yabo Wang, Rebing Wu
We propose a Trace-distance binary Tree AdaBoost (TTA) multi-class quantum classifier, a practical pipeline for quantum multi-class classification that combines quantum-aware reductions with ensemble learning to improve trainability and resource efficiency. TTA builds a hierarchical binary tree by choosing, at each…
Emanuel Casmiry, Neema Mduma, Ramadhani Sinde
In the face of increasing cyberattacks, Structured Query Language (SQL) injection remains one of the most common and damaging types of web threats, accounting for over 20% of global cyberattack costs. However, due to its dynamic and variable nature, the current detection methods often suffer from high false positive…
Osvaldo Velazquez-Gonzalez, Antonio Alarcón-Paredes, Cornelio Yañez-Marquez
Classification is a central task in machine learning, underpinning applications in domains such as finance, medicine, engineering, information technology, and biology. However, machine learning pattern classification can become a complex or even inexplicable task for current robust models due to the complexity of…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…
Sarah Nassar
—This paper summarizes the research conducted for a malware detection project using the Canadian Institute for Cybersecurity's MalMemAnalysis-2022 dataset. The purpose of the project was to explore the effectiveness and efficiency of machine learning techniques for the task of binary classification (i.e., benign or…
Authors not listed
Accurately predicting chemical reaction yields in silico is a long-standing goal in organic chemistry that, if achieved, would revolutionize synthesis design, op-timization, and discovery. The vast reaction data within scientific literature rep-resents a rich resource for training predictive machine learning models…
Marie-Luise Leitner, Martin Arendasy
Logistic regression is one of the most widely used and foundational models in both psychological research and statistical classification. As a generalized linear model (GLM), it provides a robust framework for estimating the probability of a binary outcome based on one or more predictor variables (; ). Its enduring…
Rossana O. Souza, Wellington Francisco Rodrigues, Bráulio R. G. M. Couto, Marcos A. dos Santos
Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic…
Areen Arabiat, Hamza Abu Owida, Suhaila Abuowaida, Nawaf Alshdaifat + 2 more
This study emphasizes the potential of computational techniques in cancer risk assessment, highlighting opportunities for specific and data-driven healthcare solutions. It examines the use of artificial intelligence (AI), machine learning (ML), and deep learning (DL) approaches to improve cancer risk assessment using a…
Nigmet Koklu
In recent years, evaluating competencies such as knowledge, practical skills, character traits, and meta-learning capabilities has gained increasing importance in educational research. As educational datasets grow larger and more complex, machine learning offers promising tools for analyzing student responses and…
Szymon Wojciechowski, Michał Woźniak
Many machine learning tasks aim to find models that work well not for a single, but for a group of criteria, often opposing ones. One such example is imbalanced data classification, where, on the one hand, we want to achieve the best possible classification quality for data from the minority class without degrading the…
Andrea Lo Sasso, Nicola Amoroso, Domenico Diacono, Marianna La Rocca + 6 more
Breath analysis is emerging as a non-invasive and promising diagnostic approach capable of assessing a patient’s metabolic state by detecting volatile organic compounds in exhaled breath. This study investigates the potential of breath analysis for the early detection of lung cancer, respiratory and gastrointestinal…
Julie R. Pivin-Bachler, Egon L. van den Broek
Title: Summary Ranging from health to cybersecurity, real-world data are heavily imbalanced. Handling imbalance is among the formidable challenges of machine learning (ML), as it deteriorates ML’s performance, yielding biased results toward majority classes. However, finding an adequate measure to assess the impact of…
Authors not listed
Monoterpene synthases (mTSs) are a large family of enzymes, which have promising industrial applications, yet remain difficult to engineer due to complex and poorly understood sequence-function relationships. Here, we present a structure-based machine learning (ML) framework that accurately predicts whether a mTS…
Célian Monchy, Olivier Gimenez, Céline Le Bohec, Gaël Bardon + 1 more
Deep learning (DL) is increasingly integrated into quantitative ecology, particularly for automating the classification of sensor data in biodiversity monitoring. In addition to substantially reducing data processing effort, DL models often achieve high classification performance. However, despite ongoing improvements…
Authors not listed
Quantitative Structure Activity Relationship (QSAR) remains an effective tool for early-stage chemical modelling and virtual screening in drug design. The advancements in this field are led by two core paradigms, 1) descriptor engineering, where complex fixed-length vectors of compounds are generated and conventional…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Preston Raab, W. Evan Johnson, Stephen R. Piccolo
Precision medicine relies on accurate and generalizable predictions for patients across the spectrum of human diversity. Because capturing biological heterogeneity requires large sample sizes, researchers must often aggregate data from several experimental batches or independent studies. This integration allows for…
Abdelmonem M. Ibrahim, Doaa A. Fakhry, Fares Al-Shargie, Jan Cornelis
Feature selection is crucial for high-dimensional sensor and biomedical data because it reduces redundancy, improves generalization, and supports interpretable biomarker discovery. In this study, we propose a Binary Chaos-Enhanced Newton-Raphson-Based Optimizer (BCNRBO) for wrapper-based feature selection. The method…
Maximilian Poretschkin, Tabea Naeven
The European AI Act is the first comprehensive regulation of artificial intelligence (AI), setting out extensive obligations, particularly for so-called high-risk and general-purpose AI systems. A key distinguishing feature of AI systems under the AI Act is the capability to infer. Since the AI Act does not clearly…
Leonid Sidorov, Anna Makarova, Archil Maysuradze, Mikhail Lebedev
Accurate detection of P300 event-related potentials from electroencephalography (EEG) re-mains challenging for small numbers of trials due to low signal-to-noise ratios and substantial inter-subject variability. This study presents a systematic comparison of data aggregation strate-gies for improving P300…
Authors not listed
DNA-encoded libraries (DELs) have emerged as a powerful platform for screening ultra-large chemical spaces by leveraging DNA barcodes to tag and track individual small molecules. Recent work has shown that machine learning can enhance DEL based hit discovery by denoising sequencing artifacts and improving binder…
Al Mukshit Plabon, Abdul Mukit, Md. Neyamul, Omar Faruk Jehady + 3 more
Interictal epileptiform discharges (IEDs) are diagnostically important EEG abnormalities observed between seizures. This study addresses a conditional spatial-classification task where every analyzed four-second epoch had already been reviewed and confirmed by experts as containing an IED, and the model assigned that…