23 papers · ranked by Valyu relevance
Soham Sarkar, Anil K. Ghosh
Popular clustering algorithms based on usual distance functions (e.g., Euclidean distance) often suffer in high dimension, low sample size (HDLSS) situations, where concentration of pairwise distances has adverse effects on their performance. In this article, we use a dissimilarity measure based on the data cloud…
Guoqi Qian, Yuehua Wu, Davide Ferrari, Puxue Qiao + 1 more
'Frédéric Hollande'] Regression clustering is a mixture of unsupervised and supervised statistical learning and data mining method which is found in a wide range of applications including artificial intelligence and neuroscience. It performs unsupervised learning when it clusters the data according to their respective…
Johann M Kraus, Hans A Kestler
Background In recent years, the demand for computational power in computational biology has increased due to rapidly growing data sets from microarray and other high-throughput technologies. This demand is likely to increase. Standard algorithms for analyzing data, such as cluster algorithms, need to be parallelized…
Kariyam, Abdurakhman, Adhitya Ronnie Effendie
Most existing methods of determining the number of groups apply to particular data types or are calculated based on the distance matrix for all object pairs. In this paper, we propose a medoid-based Deviation Ratio Index (DRI) to determine the number of clusters. The DRI is calculated based on the distance matrix for…
Jayasree Saha, Jayanta Mukherjee
Determining the number of clusters present in a dataset is an important problem in cluster analysis. Conventional clustering techniques generally assume this parameter to be provided up front. In this paper, we propose a method which analyzes cluster stability for predicting the cluster number. Under the same…
N. Nidheesh, K.A. Abdul Nazeer, P.M. Ameer
Cancer subtype discovery from omics data requires techniques to estimate the number of natural clusters in the data. Automatically estimating the number of clusters has been a challenging problem in Machine Learning. Using clustering algorithms together with internal cluster validity indexes have been a popular method…
Christopher N Foley, Paul D W Kirk, Stephen Burgess
Mendelian randomization is an epidemiological technique that uses genetic variants as instrumental variables to estimate the causal effect of a risk factor on an outcome. We consider a scenario in which causal estimates based on each variant in turn differ more strongly than expected by chance alone, but the variants…
Fatima Batool
A unified clustering approach that can estimate number of clusters and produce clustering against this number simultaneously is proposed. Average silhouette width (ASW) is a widely used standard cluster quality index. A distance based objective function that optimizes ASW for clustering is defined. The proposed…
Freweyni K. Teklehaymanot, Michael Muma, Abdelhak M. Zoubir
—A major challenge in cluster analysis is that the number of data clusters is mostly unknown and it must be estimated prior to clustering the observed data. In real-world applications, the observed data is often subject to heavy tailed noise and outliers which obscure the true underlying structure of the data.…
Aasim Ayaz Wani, Davide Chicco
This survey rigorously explores contemporary clustering algorithms within the machine learning paradigm, focusing on five primary methodologies: centroid-based, hierarchical, density-based, distribution-based, and graph-based clustering. Through the lens of recent innovations such as deep embedded clustering and…
Christian Hennig, Pietro Coretto
We introduce a new approach to deciding the number of clusters. The approach is applied to Optimally Tuned Robust Improper Maximum Likelihood Estimation (OTRIMLE; Coretto and Hennig (2016)) of a Gaussian mixture model allowing for observations to be classified as "noise", but it can be applied to other clustering…
S. Wade
Bayesian cluster analysis offers substantial benefits over algorithmic approaches by providing not only point estimates but also uncertainty in the clustering structure and patterns within each cluster. An overview of Bayesian cluster analysis is provided, including both model-based and loss-based approaches, along…
Joao C. Marques, Michael B. Orger
How to partition a data set into a set of distinct clusters is a ubiquitous and challenging problem. The fact that data sets vary widely in features such as cluster shape, cluster number, density distribution, background noise, outliers and degree of overlap, makes it difficult to find a single algorithm that can be…
Soumita Modak
A new clustering accuracy measure is proposed to determine the unknown number of clusters and to assess the quality of clustering of a data set given in any dimensional space. Our validity index applies the classical nonparametric univariate kernel density estimation method to the interpoint distances computed between…
Rituparna Khan, Xian Mallory
Many cancer genomes have been known to contain more than one subclone inside one tumor, the phenomenon of which is called intra-tumor heterogeneity (ITH). Characterizing ITH is essential in designing treatment plans, prognosis as well as the study of cancer progression. Single-cell DNA sequencing (scDNAseq) has been…
William Weimin Yoo
Clustering is widely studied in statistics and machine learning, with applications in a variety of fields. As opposed to popular algorithms such as agglomerative hierarchical clustering or k-means which return a single clustering solution, Bayesian nonparametric models provide a posterior over the entire space of…
Andrew R. Cohen, Paul M.B. Vitányi
For each partition of a data set into a given number of parts there is a partition such that every part is as much as possible a good model (an “algorithmic sufficient statistic”) for the data in that part. Since this can be done for every number between one and the number of data, the result is a function, the cluster…
Huichao Yan, Ting Chen, Peng Wang, Linmei Zhang + 3 more
'Yanping Bai' 'Bin Zhou'] Direction of arrival (DOA) estimation has always been a hot topic for researchers. The complex and changeable environment makes it very challenging to estimate the DOA in a small snapshot and strong noise environment. The direction-of-arrival estimation method based on compressed sensing (CS)…
A. Shivanandan, J. Unnikrishnan, A. Radenovic
Spatial aggregation or clustering of membrane proteins could be important for their functionality, e.g., in signaling, and nanoscale imaging can be used to study its origins, structure and function. Such studies require accurate characterization of clusters, both for absolute quantification and hypothesis testing. A…
Mary Pitman, David Hahn, Gary Tresadern, David Mobley
Drug discovery is accelerated with computational methods such as alchemical simulations to estimate ligand affinities. In particular, relative binding free energy (RBFE) simulations are beneficial for lead optimization. To use RBFE simulations to compare prospective ligands in silico, researchers first plan the…
Authors not listed
The screening of chemical libraries is an essential starting point in the drug discovery process. While some researchers desire a more thorough screening of drug targets against a narrower scope of molecules, it is not uncommon for diverse screening sets to be favored during early stages of drug discovery. However, a…
Sukriti Manna, Alberto Hernandez, Yunzhe Wang, Peter Lile + 2 more
The chemical and structural properties of atomically precise nanoclusters are of great interest in numerous applications, but the structures of the clusters can be computationally expensive to predict. In this work, we present the largest database of cluster structures and properties determined using ab-initio methods…
Andreas Buchgraitz Jensen, Jonas Elm
Atmospheric molecular clusters are important for the formation of new aerosol particles in the air. However, current experimental techniques are not able to yield direct insight into the cluster geometries. This implies that to date there is limited information about how accurately the applied computational methods…