25 papers · ranked by Valyu relevance
Gabriel O. Assunção, Rafael Izbicki, Marcos O. Prates
Imbalanced datasets present a significant challenge for machine learning models, often leading to biased predictions. To address this issue, data augmentation techniques are widely used in natural language processing (NLP) to generate new samples for the minority class. However, in this paper, we challenge the common…
Evgeny Burnaev, Pavel Erofeev, Artem Papanov
In many real-world binary classification tasks (e.g. detection of certain objects from images), an available dataset is imbalanced, i.e., it has much less representatives of a one class (a minor class), than of another. Generally, accurate prediction of the minor class is crucial but it's hard to achieve since there is…
Rawan S. Abdulsadig, Esther Rodriguez-Villegas
Class imbalance is a common challenge that is often faced when dealing with classification tasks aiming to detect medical events that are particularly infrequent. Apnoea is an example of such events. This challenge can however be mitigated using class rebalancing algorithms. This work investigated 10 widely used…
Annie Kim, Inkyung Jung, Duksan Ryu
Class imbalance is a major problem in classification, wherein the decision boundary is easily biased toward the majority class. A data-level solution (resampling) is one possible solution to this problem. However, several studies have shown that resampling methods can deteriorate the classification performance. This is…
Taejun Lee, Minju Kim, Sung-Phil Kim
The oddball paradigm used in P300-based brain-computer interfaces (BCIs) intrinsically poses the issue of data imbalance between target stimuli and nontarget stimuli. Data imbalance can cause overfitting problems and, consequently, poor classification performance. The purpose of this study is to improve BCI performance…
Carla Vairetti, José Luis Assadi, Sebastián Maldonado
Imbalanced classification is a well-known challenge faced by many real-world applications. This issue occurs when the distribution of the target variable is skewed, leading to a prediction bias toward the majority class. With the arrival of the Big Data era, there is a pressing need for efficient solutions to solve…
Haibin Lv, Yanhui Du, Xing Zhou, Wenkai Ni + 3 more
'Alexander Sim' 'Jinoh Kim'] With the rapid development of the Internet of Things (IoT), the frequency of attackers using botnets to control IoT devices in order to perform distributed denial-of-service attacks (DDoS) and other cyber attacks on the internet has significantly increased. In the actual attack process, the…
Asif Newaz, Farhan Shahriyar Haq
Class imbalance is a frequently occurring scenario in classification tasks. Learning from imbalanced data poses a major challenge, which has instigated a lot of research in this area. Data preprocessing using sampling techniques is a standard approach to deal with the imbalance present in the data. Since standard…
Firuz Kamalov
Imbalanced response variable distribution is a common occurrence in data science. In fields such as fraud detection, medical diagnostics, system intrusion detection and many others where abnormal behavior is rarely observed the data under study often features disproportionate target class distribution. One common way…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…
Pengyi Yang, Liang Xu, Bing B Zhou, Zili Zhang + 1 more
Background Medical and biological data are commonly with small sample size, missing values, and most importantly, imbalanced class distribution. In this study we propose a particle swarm based hybrid system for remedying the class imbalance problem in medical and biological data mining. This hybrid system combines the…
Saptarshi Bej, Narek Davtyan, Markus Wolfien, Mariam Nassar + 1 more
'Olaf Wolkenhauer'] The Synthetic Minority Oversampling TEchnique (SMOTE) is widely-used for the analysis of imbalanced datasets. It is known that SMOTE frequently overgeneralizes the minority class, leading to misclassifications for the majority class, and effecting the overall balance of the model. In this article…
Roberta Falcone, Angela Montanari, Laura Anderlucci
Matrix sketching is a recently developed data compression technique. An input matrix A is efficiently approximated with a smaller matrix B, so that B preserves most of the properties of A up to some guaranteed approximation ratio. In so doing numerical operations on big data sets become faster. Sketching algorithms…
Debaleena Datta, Pradeep Kumar Mallick, Jana Shafi, Jaeyoung Choi + 1 more
'Muhammad Fazal Ijaz'] Imbalance in hyperspectral images creates a crisis in its analysis and classification operation. Resampling techniques are utilized to minimize the data imbalance. Although only a limited number of resampling methods were explored in the previous research, a small quantity of work has been done.…
Adrian Godlewski, Krzysztof Solowiej, Patrycja Mojsak, Joanna Godzien + 6 more
Class imbalance remains a challenge in metabolomics research, where biological and technical variability can affect statistical inference and machine learning (ML) performance. Class-balancing algorithms address this issue by either increasing minority-class observations or reducing the number of majority-class…
Saptarshi Bej, Anne-Marie Galow, Robert David, Markus Wolfien + 1 more
The research landscape of single-cell and single-nuclei RNA sequencing is evolving rapidly, and one area that is enabled by this technology, is the detection of rare cells. An automated, unbiased and accurate annotation of rare subpopulations is challenging. Once rare cells are identified in one dataset, it will…
Sadam Al-Azani, Omer S. Alkhnbashi, Emad Ramadan, Motaz Alfarraj + 1 more
'Yuriy L. Orlov'] Cancer is a leading cause of death globally. The majority of cancer cases are only diagnosed in the late stages of cancer due to the use of conventional methods. This reduces the chance of survival for cancer patients. Therefore, early detection consequently followed by early diagnoses are important…
Priyanka Rana, Arcot Sowmya, Erik Meijering, Yang Song
Subcellular localisation of human proteins is essential to comprehend their functions and roles in physiological processes, which in turn helps in diagnostic and prognostic studies of pathological conditions and impacts clinical decision making. Since proteins reside at multiple locations at the same time and few…
Ayush Tripathi, Rupayan Chakraborty, Sunil Kumar Kopparapu
—Imbalance in the proportion of training samples belonging to different classes often poses performance degradation of conventional classifiers. This is primarily due to the tendency of the classifier to be biased towards the majority classes in the imbalanced dataset. In this paper, we propose a novel three step…
Lexin Chen, Ramón Alain Miranda-Quintana
DNA-Encoded Libraries allow for an efficient approach to synthesize and screen billions of small molecules against a target of interest. With more real-world binding data, this can improve training of machine learning models. However, one key challenge in DELs is the severe imbalances between the classes, in other…
Max C. Klein, Elijah Roberts
Enhanced sampling methods, such as forward flux sampling (FFS), hold a great deal of promise for accelerating stochastic simulations of nonequilibrium biochemical systems involving rare events. However, the description of the tradeoffs between simulation efficiency and error in FFS remains incomplete. We present a…
Roozbeh Valavi, Jane Elith, José J. Lahoz-Monfort, Gurutzeta Guillera-Arroita
The Random Forest (RF) algorithm is an ensemble of classification or regression trees, and is a widely used and high-performing machine learning technique. It is increasingly used for species distribution modelling (SDM). Many researchers use implementations of RF in the R programming language with default parameters…
Thomas Lynn, Julio Ottino, Richard Lueptow, Paul Umbanhowar
Cut-and-shuffle mixing is an instructive candidate system with which to assess the potential of machine learning (ML) as an approach to solve difficult mixing problems. We focus on a specific subset of cut-and-shuffle systems, the one-dimensional interval exchange transform. This class of mixing operations is well…
Henning Otto Brinkhaus, Kohulan Rajan, Achim Zielesny, Christoph Steinbeck
The development of deep learning-based optical chemical structure recognition (OCSR) systems has led to a need for datasets of chemical structure depictions. The diversity of the features in the training data is an important factor for the generation of deep learning systems that generalise well and are not overfit to…
Authors not listed
Metastable states and the conformational transitions in between them are key to understanding dynamical behaviour and function of large-scale molecular systems. By combining basic dimensionality reduction techniques with a state-of-the art approximation of the Koopman operator associated to molecular dynamics…