27 papers · ranked by Valyu relevance
Sawsan Kanj, Thomas Brüls, Stéphane Gazut
We present a new algorithm to cluster high dimensional sequence data, and its application to the field of metagenomics, which aims to reconstruct individual genomes from a mixture of genomes sampled from an environ-mental site, without any prior knowledge of reference data (genomes) or the shape of clusters. Such…
Aasim Ayaz Wani, Davide Chicco
This survey rigorously explores contemporary clustering algorithms within the machine learning paradigm, focusing on five primary methodologies: centroid-based, hierarchical, density-based, distribution-based, and graph-based clustering. Through the lens of recent innovations such as deep embedded clustering and…
Mayra Z. Rodriguez, Cesar H. Comin, Dalcimar Casanova, Odemir M. Bruno + 4 more
Many real-world systems can be studied in terms of pattern recognition tasks, so that proper use (and understanding) of machine learning methods in practical applications becomes essential. While many classification methods have been proposed, there is no consensus on which methods are more suitable for a given…
Lori Dalton, Virginia Ballarin, Marcel Brun
The development of microarray technology has enabled scientists to measure the expression of thousands of genes simultaneously, resulting in a surge of interest in several disciplines throughout biology and medicine. While data clustering has been used for decades in image processing and pattern recognition, in recent…
Mina Bagherzade Ghazvini, Miquel Sànchez-Marrè, Edgar Bahilo, Cecilio Angulo + 2 more
Operational modes of a process are described by a number of relevant features that are indicative of the state of the process. Hundreds of sensors continuously collect data in industrial systems, which shows how the relationship between different variables changes over time and identifies different modes of operation.…
Yi Wang, Yi Li, Chunhong Qiao, Xiaoyu Liu + 4 more
'Yin Yao Shugart' 'Momiao Xiong' 'Li Jin'] Clustering techniques are widely used in many applications. The goal of clustering is to identify patterns or groups of similar objects within a dataset of interest. However, many cluster methods are neither robust nor sensitive to noises and outliers in real data. In this…
T. Soni Madhulatha
Clustering is a common technique for statistical data analysis, which is used in many fields, including machine learning, data mining, pattern recognition, image analysis and bioinformatics. Clustering is the process of grouping similar objects into different groups, or more precisely, the partitioning of a data set…
Christian Hennig
Note: This paper is a chapter in the forthcoming Handbook of Cluster Analysis, Hennig et al. (2015). For definitions of basic clustering methods and some further methodology, other chapters of the Handbook are referred to. To read this version of the paper without the Handbook, some knowledge of cluster analysis…
Evangelos Karatzas, Maria Gkonta, Joana Hotova, Fotis A. Baltoumas + 4 more
Clustering is the process of grouping together different data objects based on similar properties. Clustering has applications in various case studies from several fields such as graph theory, image analysis, pattern recognition, statistics and others. Nowadays, there are numerous algorithms and tools able to generate…
Wong Hauchi, Daniil Lisik, Duy-Tai Dinh
This paper explores the critical role of data clustering in data science, emphasizing its methodologies, tools, and diverse applications. Traditional techniques, such as partitional and hierarchical clustering, are analyzed alongside advanced approaches such as data stream, density-based, graph-based, and model-based…
Eric Bair
Cluster analysis methods seek to partition a data set into homogeneous subgroups. It is useful in a wide variety of applications, including document processing and modern genetics. Conventional clustering methods are unsupervised, meaning that there is no outcome variable nor is anything known about the relationship…
Joao C. Marques, Michael B. Orger
How to partition a data set into a set of distinct clusters is a ubiquitous and challenging problem. The fact that data sets vary widely in features such as cluster shape, cluster number, density distribution, background noise, outliers and degree of overlap, makes it difficult to find a single algorithm that can be…
Michael Kern, Alexander Lex, Nils Gehlenborg, Chris R. Johnson
With ever-increasing amounts of data produced in biology research, scientists are in need of efficient data analysis methods. Cluster analysis, combined with visualization of the results, is one such method that can be used to make sense of large data volumes. At the same time, cluster analysis is known to be imperfect…
P. Ashok, G. M. Kadhar Nawaz, E. Elayaraja, V. Vadivel
Clustering is a separation of data into groups of similar objects. Every group called cluster consists of objects that are similar to one another and dissimilar to objects of other groups. In this paper, the K-Means algorithm is implemented by three distance functions and to identify the optimal distance function for…
Stijn van Dongen
Leiden and mcl are both clustering algorithms that produce flat clusterings and have a single parameter to control cluster granularity. Leiden is based on optimisation of modularity and falls within a class of modularity optimisation approaches including FastModularity and Louvain . Leiden has a resolution parameter γ…
T Soni Madhulatha
Clustering is a common technique for statistical data analysis, Clustering is the process of grouping the data into classes or clusters so that objects within a cluster have high similarity in comparison to one another, but are very dissimilar to objects in other clusters. Dissimilarities are assessed based on the…
Martin C Nwadiugwu
The current study seeks to compare 3 clustering algorithms that can be used in gene-based bioinformatics research to understand disease networks, protein-protein interaction networks, and gene expression data. Denclue, Fuzzy-C, and Balanced Iterative and Clustering using Hierarchies (BIRCH) were the 3 gene-based…
Polina Bombina, Dwayne Tally, Zachary B. Abrams, Kevin R. Coombes
Unsupervised clustering is an important task in biomedical science. We developed a new clustering method, called SillyPutty, for unsupervised clustering. As test data, we generated a series of datasets using the Umpire R package. Using these datasets, we compared SillyPutty to several existing algorithms using multiple…
Saeedeh Pourahmad, Atefeh Basirat, Amir Rahimi, Marziyeh Doostfatemeh
'Marziyeh Doostfatemeh'] Random selection of initial centroids (centers) for clusters is a fundamental defect in K-means clustering algorithm as the algorithm's performance depends on initial centroids and may end up in local optimizations. Various hybrid methods have been introduced to resolve this defect in K-means…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Yijia Li, Jonathan Nguyen, David Anastasiu, Edgar A. Arriaga
With the aim of analyzing large-sized multidimensional single-cell datasets, we are describing our method for Cosine-based Tanimoto similarity-refined graph for community detection using Leiden’s algorithm (CosTaL). As a graph-based clustering method, CosTaL transforms the cells with high-dimensional features into a…
Authors not listed
This work investigates different formulations of internal Coordinates for molecular dynamics (MD) simulations. The goal is to assess their advantages and limitations. Furthermore, a method is presented that evaluates the quality of the partitioning of molecular structural data into clusters based on statistical…
Kamen Petrov, Andreas Bender
Organizing and partitioning sets of chemical structures is of considerable practical significance e.g. in compound library analysis and the post-processing of screening hit lists. Approaches such as unsupervised clustering are computationally demanding and dataset-dependent; on the other hand, rule-based methods, such…
Mary Pitman, David Hahn, Gary Tresadern, David Mobley
Drug discovery is accelerated with computational methods such as alchemical simulations to estimate ligand affinities. In particular, relative binding free energy (RBFE) simulations are beneficial for lead optimization. To use RBFE simulations to compare prospective ligands in silico, researchers first plan the…
Authors not listed
The screening of chemical libraries is an essential starting point in the drug discovery process. While some researchers desire a more thorough screening of drug targets against a narrower scope of molecules, it is not uncommon for diverse screening sets to be favored during early stages of drug discovery. However, a…
Hang Hu, Jyothsna Padmakumar Bindu, Julia Laskin
Mass spectrometry imaging (MSI) is widely used for the label-free molecular mapping of biological samples. The identification of co-localized molecules in MSI data is crucial to the understanding of biochemical pathways. However, complex MSI data are too large for manual annotation but too small for training deep…
Authors not listed
Recent advances in artificial intelligence have significantly improved spectral data analysis. In this study, we used unsupervised machine learning to classify chemical compounds based on infrared (IR) spectral images, without relying on prior chemical knowledge. The potential of machine learning for chemical…