28 papers · ranked by Valyu relevance
Dania Dallah, Hana Sulieman, Ayman Al Zaatreh, Firuz Kamalov + 1 more
'Kichun Lee'] Outlier detection plays a key role in data analysis by improving data quality, uncovering data entry errors, and spotting unusual patterns, such as fraudulent activities. Choosing the right detection method is essential, as some approaches may be too complex or ineffective depending on the data…
Zeynel Cebeci, Cagatay Cebeci, Yalcin Tahtali, Lutfi Bayyurt + 1 more
'Yilun Shang'] Outliers are data points that significantly deviate from other data points in a data set because of different mechanisms or unusual processes. Outlier detection is one of the intensively studied research topics for identification of novelties, frauds, anomalies, deviations or exceptions in addition to…
Sunil Kumar, Sudeep Varshney, Usha Jain, Prashant Johri + 5 more
'Abdulaziz S. Almazyad' 'Ali Wagdy Mohamed' 'Mehdi Hosseinzadeh' 'Mohammad Shokouhifar' 'Razieh Sheikhpour'] Outlier detection is essential for identifying unusual patterns or observations that significantly deviate from the normal behavior of a dataset. With the rapid growth of data science, the prevalence of…
Pete R. Jones
This paper considers how best to identify statistical outliers in psychophysical datasets, where the underlying sampling distributions are unknown. Eight methods are described, and each is evaluated using Monte Carlo simulations of a typical psychophysical experiment. The best method is shown to be one based on a…
Agnieszka Nowak-Brzezińska, Weronika Łazarz, Karsten Keller
Detecting outliers is a widely studied problem in many disciplines, including statistics, data mining, and machine learning. All anomaly detection activities are aimed at identifying cases of unusual behavior compared to most observations. There are many methods to deal with this issue, which are applicable depending…
Michiel Nijhuis, Iman van Lelyveld, Gwo-Hshiung Tzeng, Sun-Weng Huang
Outliers are often present in data and many algorithms exist to find these outliers. Often we can verify these outliers to determine whether they are data errors or not. Unfortunately, checking such points is time-consuming and the underlying issues leading to the data error can change over time. An outlier detection…
Firuz Kamalov, Ho-Hon Leung
High-dimensional data poses unique challenges in outlier detection process. Most of the existing algorithms fail to properly address the issues stemming from a large number of features. In particular, outlier detection algorithms perform poorly on data set of small size with a large number of features. In this paper…
Vijendra Singh, P. Shivani
Outliers are the points which are different from or inconsistent with the rest of the data. They can be novel, new, abnormal, unusual or noisy information. Outliers are sometimes more interesting than the majority of the data. The main challenges of outlier detection with the increasing complexity, size and variety of…
Agnieszka Duraj, Daniel Duczymiński, Małgorzata Przybyła-Kasperek
The present article is devoted to outlier detection in phases of human movement. The aim was to find the most efficient machine learning method to detect abnormal segments inside physical activities in which there is a probability of origin from other activities. The problem was reduced to a classification task. The…
Juan A. Lara, Lizcano, David, Víctor Rampérez + 2 more
Outlier detection is an important problem occurring in a wide range of areas. Outliers are the outcome of fraudulent behaviour, mechanical faults, human error or simply natural deviations. Many data mining applications perform outlier detection, often as a preliminary step in order to filter out outliers and build more…
Kostas Kolomvatsos, Christos Anagnostopoulos
—The combination of the Internet of Things and the Edge Computing gives many opportunities to support innovative applications close to end users. Numerous devices present in both infrastructures can collect data upon which various processing activities can be performed. However, the quality of the outcomes may be…
Amulya Agarwal, Nitin Gupta
An outlier is an observation or a data point that is far from rest of the data points in a given dataset or we can be said that an outlier is away from the center of mass of observations. Presence of outliers can skew statistical measures and data distributions which can lead to misleading representation of the…
M. H. Marghny, Ahmed I. Taloba
In this article, we present an algorithm that provides outlier detection and data clustering simultaneously. The algorithmimprovesthe estimation of centroids of the generative distribution during the process of clustering and outlier discovery. The proposed algorithm consists of two stages. The first stage consists of…
Jiaxin Cai, Weiwei Hu, Yuhui Yang, Hong Yan + 1 more
Background Outliers, data points that significantly deviate from the norm, can have a substantial impact on statistical inference and provide valuable insights in data analysis. Multiple methods have been developed for outlier detection, however, almost all available approaches fail to consider the spatial dependence…
Devanshi Gupta, Arsh Gupta
Single-cell RNA sequencing (scRNA-seq) enables detailed analysis of cellular heterogeneity, but standard clustering approaches can miss rare cell populations with important disease relevance. We present an outlier-aware framework that complements clustering by identifying statistically unusual cells. Applied to…
Hossein Estiri, Shawn N. Murphy
Identifying implausible clinical observations (e.g., laboratory test and vital sign values) in Electronic Health Record (EHR) data using rule-based procedures is challenging. Anomaly/outlier detection methods can be applied as an alternative algorithmic approach to flagging such implausible values in EHRs. The primary…
Ch Priyanka, Vivek Vivek
In this paper, we had built the online model which are built incrementally by using online outlier detection algorithms under the streaming environment. We identified that there is highly necessity to have the streaming models to tackle the streaming data. The objective of this project is to study and analyze the…
Agnieszka Duraj, Piotr S. Szczepaniak, Artur Sadok, Gianni D’Angelo
This paper presents a comparative analysis of selected deep learning methods applied to anomaly detection in data streams. The anomaly detection results obtained on the popular Yahoo! Webscope S5 dataset are used for the computational experiments. The two commonly used and recommended models in the literature, which…
Hossein Estiri, Shawn N Murphy
To evaluate the utility of encoding for outlier detection in clinical observation data from Electronic Health Records (EHR). This article presents a semi-supervise encoding approach (super-encoding) for constructing a non-linear exemplar data distribution from EHR data and detecting non-conforming observations as…
Authors not listed
Accurate prediction of melting points for pure molecules remains a significant challenge in predictive chemistry, with implications across various scientific fields, including materials science, drug discovery, and separations chemistry. Traditional methods, such as group contribution (GC) techniques, have shown…
Eunice Carrasquinha, André Veríssimo, Susana Vinga
Survival analysis is a well known technique in the medical field. The identification of individuals whose survival time is too short or to long given their profile, assumes great importance for the detection of new prognostic factors. The study of these outlying observations have gained increasing relevancy with the…
Mert Demirarslan, Aslı Suner
In disease diagnosis classification, ensemble learning algorithms enable strong and successful models by training more than one learning function simultaneously. This study aimed to eliminate the irrelevant variable problem with the proposed new feature selection method and compare the ensemble learning algorithms’…
Chi Zhang, Dmytro Antypov, Matthew J Rosseinsky, Matthew Stephen Dyer
Machine learning has found wide application in the materials field, particularly in discovering structure-property relationships. However, its potential in predicting synthetic accessibility of materials remains relatively unexplored due to the lack of negative data. In this study, we employ several one-class…
Victor H. R. Nogueira, Rishabh Sharma, Rafael V. C. Guido, Michael J. Keiser
As efforts to improve the robustness of molecular representations advance, so does the need for methods to test and validate them. We use a Variational Auto-Encoder (VAE), an unsupervised deep learning model, to generate anomalous samples of a well-known molecular string format called SELF-referencIng Embedded Strings…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…
Van N. T. La, Stanley Nicholson, Amna Haneef, Lulu Kang + 1 more
Some data are just underappreciated. Maybe they look different or come from a different background than most other data. Maybe they don't fit neatly into common notions of what data on a ``curve'' should look like. Whatever the case, they are pigeonholed into a restricted role that limits their contributions. But if…
William McCorkindale, Mihajlo Filep, Nir London, Alpha A. Lee + 1 more
High throughput and rapid biological evaluation of small molecules is an essential factor in drug discovery and development. Direct-to-Biology (D2B), whereby compound purification is foregone, has emerged as a viable technique in time efficient screening, specifically for PROTAC design and biological evaluation.…
Authors not listed
The water solubility of organic molecules is critical for optimizing the performance and stability of aqueous flow batteries, as well as for various other applications. Although relatively straightforward to measure in some cases, the theoretical prediction of the solubility remains a considerable challenge. To this…