Search · four archives
Search · four archives
29 papers · ranked by Valyu relevance
Harsh Chhajer, Rahul Roy
Quantitative experiments are essential for investigating, uncovering, and confirming our understanding of complex systems, necessitating the use of effective and robust experimental designs. Despite generally outperforming other approaches, the broader adoption of model-based design of experiments (MBDoE) has been…
Julio-Omar Palacio-Niño, Fernando Berzal
—Determining the quality of the results obtained by clustering techniques is a key issue in unsupervised machine learning. Many authors have discussed the desirable features of good clustering algorithms. However, Jon Kleinberg established an impossibility theorem for clustering. As a consequence, a wealth of studies…
Soumita Modak
Cluster analysis is a widely applied machine learning technique to understand the existing patterns in the population of gamma-ray bursts (GRBs), in order to explore their physical sources. In the present scenario, the number of clusters corresponding to differentiable groups is still under conflict, in spite of…
Guoqi Qian, Yuehua Wu, Davide Ferrari, Puxue Qiao + 1 more
'Frédéric Hollande'] Regression clustering is a mixture of unsupervised and supervised statistical learning and data mining method which is found in a wide range of applications including artificial intelligence and neuroscience. It performs unsupervised learning when it clusters the data according to their respective…
S. Wade
Bayesian cluster analysis offers substantial benefits over algorithmic approaches by providing not only point estimates but also uncertainty in the clustering structure and patterns within each cluster. An overview of Bayesian cluster analysis is provided, including both model-based and loss-based approaches, along…
Dian Mo, Marco F. Duarte
Compressive sensing (CS) has attracted significant attention in parameter estimation tasks, where parametric dictionaries (PDs) collect signal observations for a sampling of the parameter space and yield sparse representations for signals of interest when the sampling is dense. While this sampling also leads to high…
Christopher R. John, David Watson, Dominic Russ, Katriona Goldmann + 4 more
Genome-wide data is used to stratify patients into classes for precision medicine using clustering algorithms. A common problem in this area is selection of the number of clusters (K). The Monti consensus clustering algorithm is a widely used method which uses stability selection to estimate K. However, the method has…
Aasim Ayaz Wani, Davide Chicco
This survey rigorously explores contemporary clustering algorithms within the machine learning paradigm, focusing on five primary methodologies: centroid-based, hierarchical, density-based, distribution-based, and graph-based clustering. Through the lens of recent innovations such as deep embedded clustering and…
Yasin Senbabaoglu, George Michailidis, Jun Z. Li
Consensus clustering (CC) is an unsupervised class discovery method widely used to study sample heterogeneity in high-dimensional datasets. It calculates “consensus rate” between any two samples as how frequently they are grouped together in repeated clustering runs under a certain degree of random perturbation. The…
Mayra Z. Rodriguez, Cesar H. Comin, Dalcimar Casanova, Odemir M. Bruno + 4 more
Many real-world systems can be studied in terms of pattern recognition tasks, so that proper use (and understanding) of machine learning methods in practical applications becomes essential. While many classification methods have been proposed, there is no consensus on which methods are more suitable for a given…
Paulina Pankowska, Daniel L. Oberski
Clustering consists of a popular set of techniques used to separate data into interesting groups for further analysis. Many data sources on which clustering is performed are well-known to contain random and systematic measurement errors. Such errors may adversely affect clustering. While several techniques have been…
Pietro Coretto, Christian Hennig
> Abstract. The two main topics of this paper are the introduction of the "optimally tuned improper maximum likelihood estimator" (OTRIMLE) for robust clustering based on the multivariate Gaussian model for clusters, and a comprehensive simulation study comparing the OTRIMLE to Maximum Likelihood in Gaussian mixtures…
Wanli Zhang, Yanming Di
Model-based clustering with finite mixture models has become a widely used clustering method. One of the recent implementations is MCLUST. When objects to be clustered are summary statistics, such as regression coefficient estimates, they are naturally associated with estimation errors, whose covariance matrices can…
Mina Mirshahi, Vahid Partovi-Nia, Masoud Asgharian
Shape is an important phenotype of living species that contain different environmental and genetic information. Clustering living cells using their shape information can provide a preliminary guide to their functionality and evolution. Hierarchical clustering and dendrograms, as a visualization tool for hierarchical…
Stephen Coleman, Paul D.W. Kirk, Chris Wallace
Cluster analysis is an integral part of precision medicine and systems biology, used to define groups of patients or biomolecules. However, problems such as choosing the number of clusters and issues with high dimensional data arise consistently. An ensemble approach, such as consensus clustering, can overcome some of…
Mary Pitman, David Hahn, Gary Tresadern, David Mobley
Drug discovery is accelerated with computational methods such as alchemical simulations to estimate ligand affinities. In particular, relative binding free energy (RBFE) simulations are beneficial for lead optimization. To use RBFE simulations to compare prospective ligands in silico, researchers first plan the…
Jayasree Saha, Jayanta Mukherjee
Determining the number of clusters present in a dataset is an important problem in cluster analysis. Conventional clustering techniques generally assume this parameter to be provided up front. In this paper, we propose a method which analyzes cluster stability for predicting the cluster number. Under the same…
Behnam Yousefi, Benno Schwikowski
Clustering plays an important role in a multitude of bioinformatics applications, including protein function prediction, population genetics, and gene expression analysis. The results of most clustering algorithms are sensitive to variations of the input data, the clustering algorithm and its parameters, and individual…
J W G Addy, J Langhorne
Assessing how adequate clusters fit a dataset and finding an optimum number of clusters is a difficult process. A membership matrix and the degree of membership matrix is suggested to determine the homogeneity of a cluster fit. Maximisation of the ratio of the overall degree of membership at cluster number lag 1 is…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Neo Christopher Chung
Clustering is routinely applied to modern high-dimensional data, including gene expression measurements from microarray and RNA-seq. Iteratively estimating the cluster centers and assigning memberships according to pre-defined criteria, the clustering algorithms classify genes or samples to help ascertain molecular…
Mayra Z. Rodriguez, César H. Comin, Dalcimar Casanova, Odemir Martinez Bruno + 3 more
'Odemir Martinez Bruno' 'Diego R. Amancio' 'Francisco A. Rodrigues' 'Luciano da Fontoura Costa'] Many real-world systems can be studied in terms of pattern recognition tasks, so that proper use (and understanding) of machine learning methods in practical applications becomes essential. While a myriad of classification…
Authors not listed
The screening of chemical libraries is an essential starting point in the drug discovery process. While some researchers desire a more thorough screening of drug targets against a narrower scope of molecules, it is not uncommon for diverse screening sets to be favored during early stages of drug discovery. However, a…
Evangelos Karatzas, Maria Gkonta, Joana Hotova, Fotis A. Baltoumas + 4 more
Clustering is the process of grouping together different data objects based on similar properties. Clustering has applications in various case studies from several fields such as graph theory, image analysis, pattern recognition, statistics and others. Nowadays, there are numerous algorithms and tools able to generate…
Owen Madin, Michael Shirts
Dispersion-repulsion interactions, commonly represented in atomistic force fields by the Lennard-Jones (LJ) potential, play an important role in the accuracy of molecular simulations. Training the force field parameters used in the LJ potential is challenging, generally requiring adjustment based on simulations of…
Authors not listed
Elucidating Collective Variables (CVs) for biomolecular dynamics is crucial for understanding numerous biological processes. By leveraging the tensor-train data structure, a multilinear version of the AMUSE (Algorithm for Multiple Unknown Signals) algorithm for Koopman approximation (AMUSEt) was recently developed to…
Stijn van Dongen
Although RCL is parameter-free, the choice of input ensemble (Section 8) still merits attention. A question not addressed here is whether it is beneficial to have larger input ensembles with more frequently sampled resolution or inflation parameters. A possible addition to RCL may then be to allow reduction of the…
Authors not listed
Recent advances in artificial intelligence have significantly improved spectral data analysis. In this study, we used unsupervised machine learning to classify chemical compounds based on infrared (IR) spectral images, without relying on prior chemical knowledge. The potential of machine learning for chemical…
Authors not listed
This work investigates different formulations of internal Coordinates for molecular dynamics (MD) simulations. The goal is to assess their advantages and limitations. Furthermore, a method is presented that evaluates the quality of the partitioning of molecular structural data into clusters based on statistical…