27 papers · ranked by Valyu relevance
Guillaume Lemaître, Fernando Nogueira, Christos K. Aridas
imbalanced-learn is an open-source python toolbox aiming at providing a wide range of methods to cope with the problem of imbalanced dataset frequently encountered in machine learning and pattern recognition. The implemented state-of-the-art methods can be categorized into 4 groups: (i) under-sampling, (ii)…
Huaping Guo, Weimei Zhi, Hongbing Liu, Mingliang Xu
In recent years, imbalanced learning problem has attracted more and more attentions from both academia and industry, and the problem is concerned with the performance of learning algorithms in the presence of data with severe class distribution skews. In this paper, we apply the well-known statistical model logistic…
Prabhant Singh, Joaquin Vanschoren
Automated Machine Learning has grown very successful in automating the time-consuming, iterative tasks of machine learning model development. However, current methods struggle when the data is imbalanced. Since many real-world datasets are naturally imbalanced, and improper handling of this issue can lead to quite…
Antonio Guillén-Teruel, Marcos Caracena, Jose A. Pardo, Fernando de-la-Gándara + 2 more
'Fernando de-la-Gándara' 'José Palma' 'Juan A. Botía'] This research addresses the challenges of handling unbalanced datasets for binary classification tasks. In such scenarios, standard evaluation metrics are often biased by the disproportionate representation of the minority class. Conducting experiments across seven…
Wenhao Zhang, Ramin Ramezani, Arash Naeim
Machine learning classifiers often stumble over imbalanced datasets where classes are not equally represented. This inherent bias towards the majority class may result in low accuracy in labeling minority class. Imbalanced learning is prevalent in many real-world applications, such as medical research, network…
Tianlun Zhang, Xi Yang
Imbalanced Learning is an important learning algorithm for the classification models, which have enjoyed much popularity on many applications. Typically, imbalanced learning algorithms can be partitioned into two types, i.e., data level approaches and algorithm level approaches. In this paper, the focus is to develop a…
Waleed Albattah, Rehan Ullah Khan
The exponential growth of image and video data motivates the need for practical real-time content-based searching algorithms. Features play a vital role in identifying objects within images. However, feature-based classification faces a challenge due to uneven class instance distribution. Ideally, each class should…
Julie R. Pivin-Bachler, Egon L. van den Broek
Title: Summary Ranging from health to cybersecurity, real-world data are heavily imbalanced. Handling imbalance is among the formidable challenges of machine learning (ML), as it deteriorates ML’s performance, yielding biased results toward majority classes. However, finding an adequate measure to assess the impact of…
Cristoforo Decaro, Giovanni Battista Montanari, Marco Bianconi, Gaetano Bellanca
In spite of machine learning has been successfully used in a wide range of healthcare applications, there are several parameters that could influence the performance of a machine learning system. One of the big issues for a machine learning algorithm is related to imbalanced dataset. An imbalanced dataset occurs when…
S. Maryam Hosseini, Abubakr Shafique, Morteza Babaie, H.R. Tizhoosh
In dealing with the lack of sufficient annotated data and in contrast to supervised learning, unsupervised, self-supervised, and semi-supervised domain adaptation methods are promising approaches, enabling us to transfer knowledge from rich labeled source domains to different (but related) unlabeled target domains…
Matt Clifford, Jonathan Erskine, Alexander Hepburn, Raúl Santos‐Rodríguez + 1 more
'Raúl Santos‐Rodríguez' 'Darío García-García'] Abstract. Class imbalance poses a significant challenge in classification tasks, where traditional approaches often lead to biased models and unreliable predictions. Undersampling and oversampling techniques have been commonly employed to address this issue, yet they…
Misuk Kim, Kyu-Baek Hwang, Ryan J. Urbanowicz
In numerous classification problems, class distribution is not balanced. For example, positive examples are rare in the fields of disease diagnosis and credit card fraud detection. General machine learning methods are known to be suboptimal for such imbalanced classification. One popular solution is to balance training…
Patrick Glauner, Petko Valtchev, Radu State
The underlying paradigm of big data-driven machine learning reflects the desire of deriving better conclusions from simply analyzing more data, without the necessity of looking at theory and models. Is having simply more data always helpful? In 1936, The Literary Digest collected 2.3M filled in questionnaires to…
Philipp Thölke, Yorguin Jose Mantilla Ramos, Hamza Abdelhedi, Charlotte Maschke + 10 more
Machine learning (ML) is becoming a standard tool in neuroscience and neuroimaging research. Yet, because it is such a powerful tool, the appropriate application of ML requires a sound understanding of its subtleties and limitations. In particular, applying ML to datasets with imbalanced classes, which are very common…
Keshav Sharma, Jyoti Arora, Pooja Kherwa, Zainab Alansari + 1 more
Class imbalance is a prevalent challenge in image classification tasks, where certain classes are significantly underrepresented compared to others. This imbalance often leads to biased models that perform poorly in predicting minority classes, affecting the overall performance and reliability of image classification…
Amir Reza Salehi, Majid Khedmati
In this paper, a Cluster-based Synthetic minority oversampling technique (SMOTE) Both-sampling (CSBBoost) ensemble algorithm is proposed for classifying imbalanced data. In this algorithm, a combination of over-sampling, under-sampling, and different ensemble algorithms, including Extreme Gradient Boosting (XGBoost)…
Rawan S. Abdulsadig, Esther Rodriguez-Villegas
Class imbalance is a common challenge that is often faced when dealing with classification tasks aiming to detect medical events that are particularly infrequent. Apnoea is an example of such events. This challenge can however be mitigated using class rebalancing algorithms. This work investigated 10 widely used…
Meng Chen, Yifan Liu, John Chung Tam, Ho-yin Chan + 3 more
According to the U.S. Department of Agriculture in 2018, there are more than 100 million animals used in research, education, and testing per year. Of the laboratory animals used for research, 95 percent are mice and rats as reported by the Foundation for Biomedical Research (FBR). We present here our work in…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…
Takuto Koyama, Shigeyuki Matsumoto, Hiroaki Iwata, Ryosuke Kojima + 1 more
Identifying compound-protein interactions (CPIs) is crucial for drug discovery. Because experimentally validating CPIs is often time-consuming and costly, computational approaches are expected to facilitate the process. Rapid growths of available CPI databases have accelerated the development of many machine learning…
Priyanka Rana, Arcot Sowmya, Erik Meijering, Yang Song
Subcellular localisation of human proteins is essential to comprehend their functions and roles in physiological processes, which in turn helps in diagnostic and prognostic studies of pathological conditions and impacts clinical decision making. Since proteins reside at multiple locations at the same time and few…
Yijun Liu, Qiang Huang, Huiyan Sun, Yi Chang
It is significant but challenging to explore a subset of robust biomarkers to distinguish cancer from normal samples on high-dimensional imbalanced cancer biological omics data. Although many feature selection methods addressing high dimensionality and class imbalance have been proposed, they rarely pay attention to…
James Wellnitz, Sankalp Jain, Joshua Hochuli, Travis Maxfield + 3 more
Traditional best practices for Quantitative Structure Activity Relationship (QSAR) modeling recommend dataset balancing and balanced accuracy (BA) as the key desired objective of model development. This study challenges the conventional norms by recommending the use of models with the highest positive predictive value…
Authors not listed
Predictive maintenance (PdM) can substantially reduce unplanned downtime and maintenance costs in industrial systems. In this study, we develop a proof-of-concept XGBoost-based machine learning pipeline for detecting component faults in a hydraulic test rig using the publicly available UCI Condition Monitoring of…
Angela Lopez-del Rio, Sergio Picart, Alexandre Perera-Lluna
In silico analysis of biological activity data has become an essential technique in pharmaceutical development. Specifically, the so-called proteochemometric models aim to share information between targets in machine learning ligand-target activity prediction models. However, bioactivity datasets used in…
Authors not listed
Computational toxicology plays a pivotal role in modern drug discovery and environmental risk assessment; however, the reliability of predictive models on unseen chemical scaffolds remains a critical bottleneck. Deep learning architectures, despite their prevalence, are susceptible to ’silent failures’—yielding…
Moayad Alnammi, Shengchao Liu, Spencer S Ericksen, Gene E Ananiev + 6 more
Traditional small molecule drug discovery is a time consuming and costly endeavor. High-throughput chemical screening can only assess a tiny fraction of drug-like chemical space. The strong predictive power of modern machine learning methods for virtual chemical screening enables training models on known active and…