20 papers · ranked by Valyu relevance
Kasra Babaei, Zhiyuan Chen, Tomás Maul
Outlier detection (also known as anomaly detection or deviation detection) is a process of detecting data points in which their patterns deviate significantly from others. It is common to have outliers in industry applications, which could be generated by different causes such as human error, fraudulent activities, or…
Lei Bai, Jiasheng Wang, Yu Zhou, Sotiris Kotsiantis
Outlier mining constitutes an essential aspect of modern data analytics, focusing on the identification and interpretation of anomalous observations. Conventional density-based local outlier detection methodologies frequently exhibit limitations due to their inherent lack of data preprocessing capabilities…
Agnieszka Nowak-Brzezińska, Czesław Horyń
The article presents both methods of clustering and outlier detection in complex data, such as rule-based knowledge bases. What distinguishes this work from others is, first, the application of clustering algorithms to rules in domain knowledge bases, and secondly, the use of outlier detection algorithms to detect…
Clémentine Barreyre, Béatrice Laurent, Jean-Michel Loubès, Bertrand Cabon + 1 more
'Bertrand Cabon' 'Loïc Boussouf'] We propose a novel procedure for outlier detection in functional data, in a semisupervised framework. As the data is functional, we consider the coefficients obtained after projecting the observations onto orthonormal bases (wavelet, PCA). A multiple testing procedure based on the…
Mahmoud G. Ismail, Mohammed A.-M. Salem, Mohamed A. Abd El Ghany, Eman Abdullah Aldakheel + 2 more
'Eman Abdullah Aldakheel' 'Safia Abbas' 'Yue Zhang'] User authentication is a fundamental aspect of information security, requiring robust measures against identity fraud and data breaches. In the domain of keystroke dynamics research, a significant challenge lies in the reliance on imposter datasets, particularly…
David Muhr, Michael Affenzeller, Josef Küng
The scores of distance-based outlier detection methods are difficult to interpret, making it challenging to determine a cut-off threshold between normal and outlier data points without additional context. We describe a generic transformation of distance-based outlier scores into interpretable, probabilistic estimates.…
Zekun Xu, Deovrat Kakde, Arin Chaudhuri
In recent years, there have been many practical applications of anomaly detection such as in predictive maintenance, detection of credit fraud, network intrusion, and system failure. The goal of anomaly detection is to identify in the test data anomalous behaviors that are either rare or unseen in the training data.…
Kangqing Yu, Wei Shi, Nicola Santoro
To design an algorithm for detecting outliers over streaming data has become an important task in many common applications, arising in areas such as fraud detections, network analysis, environment monitoring and so forth. Due to the fact that real-time data may arrive in the form of streams rather than batches…
Agnieszka Nowak-Brzezińska, Igor Gaibei, Gergely Palla
In this article, we evaluate the efficiency and performance of two clustering algorithms: $AHC$ (Agglomerative Hierarchical Clustering) and $K-Means$. We are aware that there are various linkage options and distance measures that influence the clustering results. We assess the quality of clustering using the…
Zihao Li, Liumei Zhang, Geert Verdoolaege
Outlier detection is an important task in the field of data mining and a highly active area of research in machine learning. In industrial automation, datasets are often high-dimensional, meaning an effort to study all dimensions directly leads to data sparsity, thus causing outliers to be masked by noise effects in…
Clémence Allietta, Jean-Philippe Condomines, Jean–Yves Tourneret, Emmanuel Lochin
'Emmanuel Lochin'] Hyperbolic geometry has recently garnered considerable attention in machine learning due to its capacity to embed hierarchical graph structures with low distortions for further downstream processing. This paper introduces a simple framework to detect local outliers for datasets grounded in hyperbolic…
Rui Hu, Luc, Chen Chen, Yiwei Wang
—The nature of modern data is increasingly real-time, making outlier detection crucial in any data-related field, such as finance for fraud detection and healthcare for monitoring patient vitals. Traditional outlier detection methods, such as the Local Outlier Factor (LOF) algorithm, struggle with realtime data due to…
Chi Zhang, Dmytro Antypov, Matthew J Rosseinsky, Matthew Stephen Dyer
Machine learning has found wide application in the materials field, particularly in discovering structure-property relationships. However, its potential in predicting synthetic accessibility of materials remains relatively unexplored due to the lack of negative data. In this study, we employ several one-class…
Quentin Rougemont, Amanda Xuereb, Xavier Dallaire, Jean-Sébastien Moore + 10 more
Inferring the genomic basis of local adaptation is a long-standing goal of evolutionary biology. Beyond its fundamental evolutionary implications, such knowledge can guide conservation decisions for populations of conservation and management concern. Here, we investigated the genomic basis of local adaptation in the…
Margaux-Alison Fustier, Natalia E. Martínez-Ainsworth, Jonás A. Aguirre-Liguori, Anthony Venon + 15 more
In plants, local adaptation across species range is frequent. Yet, much has to be discovered on its environmental drivers, the underlying functional traits and their molecular determinants. Genome scans are popular to uncover outlier loci potentially involved in the genetic architecture of local adaptation, however…
Authors not listed
Accurate prediction of melting points for pure molecules remains a significant challenge in predictive chemistry, with implications across various scientific fields, including materials science, drug discovery, and separations chemistry. Traditional methods, such as group contribution (GC) techniques, have shown…
Petri Kemppainen, Frédéric Guillaume
Understanding the genetic basis of adaptive evolution is central to predicting how populations respond to environmental change. Genome scans and genotype–environment association methods are widely used to detect loci under selection, but because they often ignore correlations among loci (linkage disequilibrium, LD)…
Helena Martins, Kevin Caye, Keurcien Luu, Michael G.B. Blum + 1 more
Finding genetic signatures of local adaptation is of great interest for many population genetic studies. Common approaches to sorting selective loci from their genomic background focus on the extreme values of the fixation index, F_ST_, across loci. However, the computation of the fixation index becomes challenging…
Robert Verity, Caitlin Collins, Daren C. Card, Sara M. Schaal + 2 more
Genome scans are widely used to identify “outliers” in genomic data: loci with different patterns compared with the rest of the genome due to the action of selection or other non-adaptive forces of evolution. These genomic datasets are often high-dimensional, with complex correlation structures among variables, making…
Authors not listed
Due to their position-dependent admixture of the exact-exchange (EXX) energy density, local hybrid functionals (LHs) enable a flexible balance between reduced self-interaction errors and smaller static-correlation errors, allowing an escape from the usual zero-sum game between these two central aspects of the…