27 papers · ranked by Valyu relevance
Shujuan Wang, Yuntao Dai, Jihong Shen, Jingxue Xuan
With the development of artificial intelligence, big data classification technology provides the advantageous help for the medicine auxiliary diagnosis research. While due to the different conditions in the different sample collection, the medical big data is often imbalanced. The class-imbalance problem has been…
Der-Chiang Li, Qi-Shi Shi, Yao-San Lin, Liang-Sian Lin + 1 more
'Gholamreza Anbarjafari'] Oversampling is the most popular data preprocessing technique. It makes traditional classifiers available for learning from imbalanced data. Through an overall review of oversampling techniques (oversamplers), we find that some of them can be regarded as danger-information-based oversamplers…
Asif Newaz, Shahriar Hassan, Farhan Shahriyar Haq
Learning from imbalanced data is a challenging task. Standard classification algorithms tend to perform poorly when trained on imbalanced data. Some special strategies need to be adopted, either by modifying the data distribution or by redesigning the underlying classification algorithm to achieve desirable…
Damien Dablain, Bartosz Krawczyk, Nitesh V. Chawla
Machine learning (ML) is playing an increasingly important role in rendering decisions that affect a broad range of groups in society. ML models inform decisions in criminal justice, the extension of credit in banking, and the hiring practices of corporations. This posits the requirement of model fairness, which holds…
Vinod Kumar, Gotam Singh Lalotra, Ponnusamy Sasikala, Dharmendra Singh Rajput + 6 more
'Dharmendra Singh Rajput' 'Rajesh Kaluri' 'Kuruva Lakshmanna' 'Mohammad Shorfuzzaman' 'Abdulmajeed Alsufyani' 'Mueen Uddin' 'Andrea Tittarelli'] Nowadays, healthcare is the prime need of every human being in the world, and clinical datasets play an important role in developing an intelligent healthcare system for…
Misuk Kim, Kyu-Baek Hwang, Ryan J. Urbanowicz
In numerous classification problems, class distribution is not balanced. For example, positive examples are rare in the fields of disease diagnosis and credit card fraud detection. General machine learning methods are known to be suboptimal for such imbalanced classification. One popular solution is to balance training…
Philipp Thölke, Yorguin Jose Mantilla Ramos, Hamza Abdelhedi, Charlotte Maschke + 10 more
Machine learning (ML) is becoming a standard tool in neuroscience and neuroimaging research. Yet, because it is such a powerful tool, the appropriate application of ML requires a sound understanding of its subtleties and limitations. In particular, applying ML to datasets with imbalanced classes, which are very common…
Ahmad B. Hassanat, Ahmad S. Tarawneh, Ghada A. Altarawneh
For the last two decades, oversampling has been employed to overcome the challenge of learning from imbalanced datasets. Many approaches to solving this challenge have been offered in the literature. Oversampling, on the other hand, is a concern. That is, models trained on fictitious data may fail spectacularly when…
S. Maryam Hosseini, Abubakr Shafique, Morteza Babaie, H.R. Tizhoosh
In dealing with the lack of sufficient annotated data and in contrast to supervised learning, unsupervised, self-supervised, and semi-supervised domain adaptation methods are promising approaches, enabling us to transfer knowledge from rich labeled source domains to different (but related) unlabeled target domains…
Elaheh Jafarigol, Theodore B. Trafali̇s
For over two decades, detecting rare events has been a challenging task among researchers in the data mining and machine learning domain. Real-life problems inspire researchers to navigate and further improve data processing and algorithmic approaches to achieve effective and computationally efficient methods for…
Julie R. Pivin-Bachler, Egon L. van den Broek
Title: Summary Ranging from health to cybersecurity, real-world data are heavily imbalanced. Handling imbalance is among the formidable challenges of machine learning (ML), as it deteriorates ML’s performance, yielding biased results toward majority classes. However, finding an adequate measure to assess the impact of…
Amir Reza Salehi, Majid Khedmati
In this paper, a Cluster-based Synthetic minority oversampling technique (SMOTE) Both-sampling (CSBBoost) ensemble algorithm is proposed for classifying imbalanced data. In this algorithm, a combination of over-sampling, under-sampling, and different ensemble algorithms, including Extreme Gradient Boosting (XGBoost)…
Shraddha M. Naik, Tanujit Chakraborty, Abdenour Hadid, Bibhas Chakraborty
'Bibhas Chakraborty'] Real-world datasets often exhibit imbalanced data distribution, where certain class levels are severely underrepresented. In such cases, traditional pattern classifiers have shown a bias towards the majority class, impeding accurate predictions for the minority class. This paper introduces an…
Debaleena Datta, Pradeep Kumar Mallick, Jana Shafi, Jaeyoung Choi + 1 more
'Muhammad Fazal Ijaz'] Imbalance in hyperspectral images creates a crisis in its analysis and classification operation. Resampling techniques are utilized to minimize the data imbalance. Although only a limited number of resampling methods were explored in the previous research, a small quantity of work has been done.…
Annie Kim, Inkyung Jung, Duksan Ryu
Class imbalance is a major problem in classification, wherein the decision boundary is easily biased toward the majority class. A data-level solution (resampling) is one possible solution to this problem. However, several studies have shown that resampling methods can deteriorate the classification performance. This is…
Xinyi Gao, Dongting Xie, Yihang Zhang, Zhengren Wang + 3 more
'Hongzhi Yin' 'Wentao Zhang'] Abstract—With the expansion of data availability, machine learning (ML) has achieved remarkable breakthroughs in bot h academia and industry. However, imbalanced data distributions are prevalent in various types of raw data and severely hinde r the performance of ML by biasing the…
Husam Abdulnabi, J. Timothy Westwood
Machine Learning (ML) models may perform inconsistently on individual classes on nominal outputs or ranges on continuous outputs, collectively referred to here as bins. Models should be assessed through metrics that consider each bin individually, called bin metrics. Inconsistent model performance is often due to model…
Husam Abdulnabi, J. Timothy Westwood
Machine Learning (ML) models may perform inconsistently on individual classes on nominal outputs or ranges on continuous outputs, collectively referred to here as bins. Models should be assessed through metrics that consider each bin individually, called bin metrics. Inconsistent model performance is often due to model…
Mahabubur Rahman Miraj, Hongyu Huang, Ting Yang, Jiafei Zhao + 2 more
Imbalanced classification is a significant challenge in machine learning, especially in critical applications like medical diagnosis, fraud detection, and cybersecurity. Traditional oversampling techniques, such as SMOTE, often fail to handle label noise and complex data distributions, leading to reduced classification…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…
Priyanka Rana, Arcot Sowmya, Erik Meijering, Yang Song
Due to progression in cell-cycle or duration of storage, classification of morphological changes in human blood cells is important for correct and effective clinical decisions. Automated classification systems help avoid subjective outcomes and are more efficient. Deep learning and more specifically Convolutional…
Priyanka Rana, Arcot Sowmya, Erik Meijering, Yang Song
Subcellular localisation of human proteins is essential to comprehend their functions and roles in physiological processes, which in turn helps in diagnostic and prognostic studies of pathological conditions and impacts clinical decision making. Since proteins reside at multiple locations at the same time and few…
Takuto Koyama, Shigeyuki Matsumoto, Hiroaki Iwata, Ryosuke Kojima + 1 more
Identifying compound-protein interactions (CPIs) is crucial for drug discovery. Because experimentally validating CPIs is often time-consuming and costly, computational approaches are expected to facilitate the process. Rapid growths of available CPI databases have accelerated the development of many machine learning…
Authors not listed
Predictive maintenance (PdM) can substantially reduce unplanned downtime and maintenance costs in industrial systems. In this study, we develop a proof-of-concept XGBoost-based machine learning pipeline for detecting component faults in a hydraulic test rig using the publicly available UCI Condition Monitoring of…
Joseph Davies, David Pattison, Jonathan Hirst
Machine learning models were developed to predict product formation from time-series reaction data for ten Buchwald-Hartwig coupling reactions. The data was provided by DeepMatter and was collected in their DigitalGlassware cloud platform. The reaction probe has 12 sensors to measure properties of interest, including…
James Wellnitz, Sankalp Jain, Joshua Hochuli, Travis Maxfield + 3 more
Traditional best practices for Quantitative Structure Activity Relationship (QSAR) modeling recommend dataset balancing and balanced accuracy (BA) as the key desired objective of model development. This study challenges the conventional norms by recommending the use of models with the highest positive predictive value…
Authors not listed
Machine learning holds significant promise for accelerating biomarker discovery in clinical proteomics, yet its real-world impact remains limited by widespread methodological pitfalls and unrealistic expectations. In this perspective, we critically examine the integration of machine learning into clinical proteomics…