29 papers · ranked by Valyu relevance
Rushi Longadge, Snehalata Dongre
In last few years there are major changes and evolution has been done on classification of data. As the application area of technology is increases the size of data also increases. Classification of data becomes difficult because of unbounded size and imbalance nature of data. Class imbalance problem become greatest…
Shujuan Wang, Yuntao Dai, Jihong Shen, Jingxue Xuan
With the development of artificial intelligence, big data classification technology provides the advantageous help for the medicine auxiliary diagnosis research. While due to the different conditions in the different sample collection, the medical big data is often imbalanced. The class-imbalance problem has been…
Khan Md. Hasib, Md. Sadiq Iqbal, Faisal Muhammad Shah, Jubayer Al Mahmud + 4 more
'Jubayer Al Mahmud' 'Mahmudul Hasan Popel' 'Md. Imran Hossain Showrov' 'Shakil Ahmed' 'Obaidur Rahman'] Abstract: The problem of class imbalance is extensive for focusing on numerous applications in the real world. In such a situation, nearly all of the examples are labeled as one class called majority class, while far…
Misuk Kim, Kyu-Baek Hwang, Ryan J. Urbanowicz
In numerous classification problems, class distribution is not balanced. For example, positive examples are rare in the fields of disease diagnosis and credit card fraud detection. General machine learning methods are known to be suboptimal for such imbalanced classification. One popular solution is to balance training…
Vinod Kumar, Gotam Singh Lalotra, Ponnusamy Sasikala, Dharmendra Singh Rajput + 6 more
'Dharmendra Singh Rajput' 'Rajesh Kaluri' 'Kuruva Lakshmanna' 'Mohammad Shorfuzzaman' 'Abdulmajeed Alsufyani' 'Mueen Uddin' 'Andrea Tittarelli'] Nowadays, healthcare is the prime need of every human being in the world, and clinical datasets play an important role in developing an intelligent healthcare system for…
Der-Chiang Li, Qi-Shi Shi, Yao-San Lin, Liang-Sian Lin + 1 more
'Gholamreza Anbarjafari'] Oversampling is the most popular data preprocessing technique. It makes traditional classifiers available for learning from imbalanced data. Through an overall review of oversampling techniques (oversamplers), we find that some of them can be regarded as danger-information-based oversamplers…
Ajay Kulkarni, Deri Chong, Feras A. Batarseh
Dealing with imbalanced data is a prevalent problem while performing classification on the datasets. Many times, this problem contributes to bias while making decisions or implementing policies. Thus, it is vital to understand the factors which causes imbalance in the data (or class imbalance). Such hidden biases and…
Philipp Thölke, Yorguin Jose Mantilla Ramos, Hamza Abdelhedi, Charlotte Maschke + 10 more
Machine learning (ML) is becoming a standard tool in neuroscience and neuroimaging research. Yet, because it is such a powerful tool, the appropriate application of ML requires a sound understanding of its subtleties and limitations. In particular, applying ML to datasets with imbalanced classes, which are very common…
Richmond Addo Danquah
For several years till date, the major issues in terms of solving for classification problems are the issues of Imbalanced data. Because majority of the machine learning algorithms by default assumes all data are balanced, the algorithms do not take into consideration the distribution of the data sample class. The…
S. Maryam Hosseini, Abubakr Shafique, Morteza Babaie, H.R. Tizhoosh
In dealing with the lack of sufficient annotated data and in contrast to supervised learning, unsupervised, self-supervised, and semi-supervised domain adaptation methods are promising approaches, enabling us to transfer knowledge from rich labeled source domains to different (but related) unlabeled target domains…
Huaping Guo, Weimei Zhi, Hongbing Liu, Mingliang Xu
In recent years, imbalanced learning problem has attracted more and more attentions from both academia and industry, and the problem is concerned with the performance of learning algorithms in the presence of data with severe class distribution skews. In this paper, we apply the well-known statistical model logistic…
Soroush Saryazdi, Bahareh Nikpour, Hossein Nezamabadi‐pour
—Learning from many real-world datasets is limited by a problem called the class imbalance problem. A dataset is imbalanced when one class (the majority class) has significantly more samples than the other class (the minority class). Such datasets cause typical machine learning algorithms to perform poorly on the…
Amir Reza Salehi, Majid Khedmati
In this paper, a Cluster-based Synthetic minority oversampling technique (SMOTE) Both-sampling (CSBBoost) ensemble algorithm is proposed for classifying imbalanced data. In this algorithm, a combination of over-sampling, under-sampling, and different ensemble algorithms, including Extreme Gradient Boosting (XGBoost)…
Mahmoud El-Banna
The Mahalanobis Taguchi System (MTS) is considered one of the most promising binary classification algorithms to handle imbalance data. Unfortunately, MTS lacks a method for determining an efficient threshold for the binary classification. In this paper, a nonlinear optimization model is formulated based on minimizing…
Damien Dablain, Bartosz Krawczyk, Nitesh V. Chawla
Machine learning (ML) is playing an increasingly important role in rendering decisions that affect a broad range of groups in society. ML models inform decisions in criminal justice, the extension of credit in banking, and the hiring practices of corporations. This posits the requirement of model fairness, which holds…
Shahzad Ashraf, Sehrish Saleem, Tauqeer Ahmed, Zeeshan Aslam + 1 more
'Durr Muhammad'] An imbalanced dataset is commonly found in at least one class, which are typically exceeded by the other ones. A machine learning algorithm (classifier) trained with an imbalanced dataset predicts the majority class (frequently occurring) more than the other minority classes (rarely occurring).…
Naman Deep Singh, Abhinav Dhall
A learning classifier must outperform a trivial solution, in case of imbalanced data, this condition usually does not hold true. To overcome this problem, we propose a novel data level resampling method - Clustering Based Oversampling for improved learning from class imbalanced datasets. The essential idea behind the…
Husam Abdulnabi, J. Timothy Westwood
Machine Learning (ML) models may perform inconsistently on individual classes on nominal outputs or ranges on continuous outputs, collectively referred to here as bins. Models should be assessed through metrics that consider each bin individually, called bin metrics. Inconsistent model performance is often due to model…
Husam Abdulnabi, J. Timothy Westwood
Machine Learning (ML) models may perform inconsistently on individual classes on nominal outputs or ranges on continuous outputs, collectively referred to here as bins. Models should be assessed through metrics that consider each bin individually, called bin metrics. Inconsistent model performance is often due to model…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…
Asif Newaz, Shahriar Hassan, Farhan Shahriyar Haq
Learning from imbalanced data is a challenging task. Standard classification algorithms tend to perform poorly when trained on imbalanced data. Some special strategies need to be adopted, either by modifying the data distribution or by redesigning the underlying classification algorithm to achieve desirable…
Priyanka Rana, Arcot Sowmya, Erik Meijering, Yang Song
Due to progression in cell-cycle or duration of storage, classification of morphological changes in human blood cells is important for correct and effective clinical decisions. Automated classification systems help avoid subjective outcomes and are more efficient. Deep learning and more specifically Convolutional…
Angela Lopez-del Rio, Sergio Picart, Alexandre Perera-Lluna
In silico analysis of biological activity data has become an essential technique in pharmaceutical development. Specifically, the so-called proteochemometric models aim to share information between targets in machine learning ligand-target activity prediction models. However, bioactivity datasets used in…
Takuto Koyama, Shigeyuki Matsumoto, Hiroaki Iwata, Ryosuke Kojima + 1 more
Identifying compound-protein interactions (CPIs) is crucial for drug discovery. Because experimentally validating CPIs is often time-consuming and costly, computational approaches are expected to facilitate the process. Rapid growths of available CPI databases have accelerated the development of many machine learning…
Yijun Liu, Qiang Huang, Huiyan Sun, Yi Chang
It is significant but challenging to explore a subset of robust biomarkers to distinguish cancer from normal samples on high-dimensional imbalanced cancer biological omics data. Although many feature selection methods addressing high dimensionality and class imbalance have been proposed, they rarely pay attention to…
Authors not listed
Predictive maintenance (PdM) can substantially reduce unplanned downtime and maintenance costs in industrial systems. In this study, we develop a proof-of-concept XGBoost-based machine learning pipeline for detecting component faults in a hydraulic test rig using the publicly available UCI Condition Monitoring of…
James Wellnitz, Sankalp Jain, Joshua Hochuli, Travis Maxfield + 3 more
Traditional best practices for Quantitative Structure Activity Relationship (QSAR) modeling recommend dataset balancing and balanced accuracy (BA) as the key desired objective of model development. This study challenges the conventional norms by recommending the use of models with the highest positive predictive value…
Mert Demirarslan, Aslı Suner
In disease diagnosis classification, ensemble learning algorithms enable strong and successful models by training more than one learning function simultaneously. This study aimed to eliminate the irrelevant variable problem with the proposed new feature selection method and compare the ensemble learning algorithms’…
Authors not listed
Machine learning holds significant promise for accelerating biomarker discovery in clinical proteomics, yet its real-world impact remains limited by widespread methodological pitfalls and unrealistic expectations. In this perspective, we critically examine the integration of machine learning into clinical proteomics…