26 papers · ranked by Valyu relevance
Mina Bagherzade Ghazvini, Miquel Sànchez-Marrè, Edgar Bahilo, Cecilio Angulo + 2 more
Operational modes of a process are described by a number of relevant features that are indicative of the state of the process. Hundreds of sensors continuously collect data in industrial systems, which shows how the relationship between different variables changes over time and identifies different modes of operation.…
Jan-Oliver Felix Kapp-Joswig, Bettina G. Keller
C lustering1 is an analytic process that identifies associations of some kind (i.e. clusters) among a number of considered objects. Phrasing this differently, 'clustering is a synonym for the decomposition of a set of entities into natural groups'.[1] The word 'natural' general appearing in this quote indicates already…
Christian Hennig
Nine popular clustering methods are applied to 42 real data sets. The aim is to give a detailed characterisation of the methods by means of several cluster validation indexes that measure various individual aspects of the resulting clusters such as small within-cluster distances, separation of clusters, closeness to a…
Cyril Esnault, Melissa Rollot, Pauline Guilmin, Jean-Daniel Zucker
The exploration of heath data by clustering algorithms allows to better describe the populations of interest by seeking the sub-profiles that compose it. This therefore reinforces medical knowledge, whether it is about a disease or a targeted population in real life. Nevertheless, contrary to the so-called conventional…
Aasim Ayaz Wani, Davide Chicco
This survey rigorously explores contemporary clustering algorithms within the machine learning paradigm, focusing on five primary methodologies: centroid-based, hierarchical, density-based, distribution-based, and graph-based clustering. Through the lens of recent innovations such as deep embedded clustering and…
Manoharan Premkumar, Garima Sinha, Manjula Devi Ramasamy, Santhoshini Sahu + 4 more
'Santhoshini Sahu' 'Chithirala Bala Subramanyam' 'Ravichandran Sowmya' 'Laith Abualigah' 'Bizuwork Derebew'] This study presents the K-means clustering-based grey wolf optimizer, a new algorithm intended to improve the optimization capabilities of the conventional grey wolf optimizer in order to address the problem of…
Stijn van Dongen
Clustering (a large class of methods) is a standard and often used approach in data analysis for separating data into groups, called clusters, often in large-scale high-dimensional data. A clustering (a data structure) is a partitioning of the data into disjoint clusters. This is sometimes called a flat clustering to…
Wong Hauchi, Daniil Lisik, Duy-Tai Dinh
This paper explores the critical role of data clustering in data science, emphasizing its methodologies, tools, and diverse applications. Traditional techniques, such as partitional and hierarchical clustering, are analyzed alongside advanced approaches such as data stream, density-based, graph-based, and model-based…
Giuseppe Agapito, Marianna Milano, Mario Cannataro, Quan Zou
Gene expression and SNPs data hold great potential for a new understanding of disease prognosis, drug sensitivity, and toxicity evaluations. Cluster analysis is used to analyze data that do not contain any specific subgroups. The goal is to use the data itself to recognize meaningful and informative subgroups. In…
Hui Yin, Amir Aryani, Stephen Petrie, Aishwarya Nambissan + 2 more
'Aland Astudillo' 'Shengyuan Cao'] Clustering algorithms aim to organize data into groups or clusters based on the inherent patterns and similarities within the data. They play an important role in today's life, such as in marketing and e-commerce, healthcare, data organization and analysis, and social media. Numerous…
M.A. Naser, Ahmed Z. Naser
Neighbors Exploration for Clustering Problems Authors: ['M.A. Naser' 'Ahmed Z. Naser'] This paper presents a novel clustering algorithm from the SPINEX (Similarity-based Predictions with Explainable Neighbors Exploration) algorithmic family. The newly proposed clustering variant leverages the concept of similarity and…
Agnieszka Nowak-Brzezińska, Igor Gaibei, Gergely Palla
In this article, we evaluate the efficiency and performance of two clustering algorithms: $AHC$ (Agglomerative Hierarchical Clustering) and $K-Means$. We are aware that there are various linkage options and distance measures that influence the clustering results. We assess the quality of clustering using the…
Polina Bombina, Dwayne Tally, Zachary B. Abrams, Kevin R. Coombes
Unsupervised clustering is an important task in biomedical science. We developed a new clustering method, called SillyPutty, for unsupervised clustering. As test data, we generated a series of datasets using the Umpire R package. Using these datasets, we compared SillyPutty to several existing algorithms using multiple…
Suruchi Jai Kumar Ahuja
A major objective of clustering is to identify groups in the data that maximizes the similarity between objects within the same cluster and minimizes the similarity between different clusters. A challenge for data clustering, and unsupervised learning in general, is that there is often no mechanism for feature…
Nassir Mohammad
A computational theory for clustering and a semi-supervised clustering algorithm is presented. Clustering is defined to be the obtainment of groupings of data such that each group contains no anomalies with respect to a chosen grouping principle and measure; all other examples are considered to be fringe points…
David P. Hofmeyr
A novel and intuitive nearest neighbours based clustering algorithm is introduced, in which a cluster is defined in terms of an equilibrium condition which balances its size and cohesiveness. The formulation of the equilibrium condition allows for a quantification of the strength of alignment of each point to a…
Abiodun M. Ikotun, Absalom E. Ezugwu, Vincent Yu
Kmeans clustering algorithm is an iterative unsupervised learning algorithm that tries to partition the given dataset into k pre-defined distinct non-overlapping clusters where each data point belongs to only one group. However, its performance is affected by its sensitivity to the initial cluster centroids with the…
Behnam Yousefi, Benno Schwikowski
Clustering plays an important role in a multitude of bioinformatics applications, including protein function prediction, population genetics, and gene expression analysis. The results of most clustering algorithms are sensitive to variations of the input data, the clustering algorithm and its parameters, and individual…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Yijia Li, Jonathan Nguyen, David Anastasiu, Edgar A. Arriaga
With the aim of analyzing large-sized multidimensional single-cell datasets, we are describing our method for Cosine-based Tanimoto similarity-refined graph for community detection using Leiden’s algorithm (CosTaL). As a graph-based clustering method, CosTaL transforms the cells with high-dimensional features into a…
Elijah Willie, Pengyi Yang, Ellis Patrick
Highly multiplexed in situ imaging cytometry assays have enabled researchers to scrutinize cellular systems at an unprecedented level. With the capability of these assays to simultaneously profile the spatial distribution and molecular features of many cells, unsupervised machine learning, and in particular clustering…
Kamen Petrov, Andreas Bender
Organizing and partitioning sets of chemical structures is of considerable practical significance e.g. in compound library analysis and the post-processing of screening hit lists. Approaches such as unsupervised clustering are computationally demanding and dataset-dependent; on the other hand, rule-based methods, such…
Authors not listed
This work investigates different formulations of internal Coordinates for molecular dynamics (MD) simulations. The goal is to assess their advantages and limitations. Furthermore, a method is presented that evaluates the quality of the partitioning of molecular structural data into clusters based on statistical…
Mary Pitman, David Hahn, Gary Tresadern, David Mobley
Drug discovery is accelerated with computational methods such as alchemical simulations to estimate ligand affinities. In particular, relative binding free energy (RBFE) simulations are beneficial for lead optimization. To use RBFE simulations to compare prospective ligands in silico, researchers first plan the…
Authors not listed
The screening of chemical libraries is an essential starting point in the drug discovery process. While some researchers desire a more thorough screening of drug targets against a narrower scope of molecules, it is not uncommon for diverse screening sets to be favored during early stages of drug discovery. However, a…
Authors not listed
Recent advances in artificial intelligence have significantly improved spectral data analysis. In this study, we used unsupervised machine learning to classify chemical compounds based on infrared (IR) spectral images, without relying on prior chemical knowledge. The potential of machine learning for chemical…