26 papers · ranked by Valyu relevance
Stephen Swift, Allan Tucker, Veronica Vinciotti, Nigel Martin + 3 more
Consensus clustering, a new method for analyzing microarray data that takes a consensus set of clusters from various algorithms, is shown to perform better than individual methods alone.
Raffaele Giancarlo, Filippo Utro
Background The inference of the number of clusters in a dataset, a fundamental problem in Statistics, Data Analysis and Classification, is usually addressed via internal validation measures. The stated problem is quite difficult, in particular for microarrays, since the inferred prediction must be sensible enough to…
Stijn van Dongen
Consensus clustering concerns the integration of multiple clusterings of the same dataset into a single result clustering , often called the consensus clustering or consensus partition. The same setting is described by many authors as the cluster(ing) ensemble problem . In this and the majority of previous work cited…
T Ian Simpson, J Douglas Armstrong, Andrew P Jarman
Background One of the most commonly performed tasks when analysing high throughput gene expression data is to use clustering methods to classify the data into groups. There are a large number of methods available to perform clustering, but it is often unclear which method is best suited to the data and how to quantify…
Awad A. Alyousef, Svetlana Nihtyanova, Chris Denton, Pietro Bosoni + 2 more
Disease subtyping, which helps to develop personalized treatments, remains a challenge in data analysis because of the many different ways to group patients based upon their data. However, if we can identify subclasses of disease, then it will help to develop better models that are more specific to individuals and…
Hongfu Liu, Zhiqiang Tao, Zhengming Ding
Consensus clustering fuses diverse basic partitions (i.e., clustering results obtained from conventional clustering methods) into an integrated one, which has attracted increasing attention in both academic and industrial areas due to its robust and effective performance. Tremendous research efforts have been made to…
Mansaf Alam, Kishwar Sadaf
Clustering of web search result document has emerged as a promising tool for improving retrieval performance of an Information Retrieval (IR) system. Search results often plagued by problems like synonymy, polysemy, high volume etc. Clustering other than resolving these problems also provides the user the easiness to…
Yunli Wang, Youlian Pan
Background Simple clustering methods such as hierarchical clustering and k-means are widely used for gene expression data analysis; but they are unable to deal with noise and high dimensionality associated with the microarray gene expression data. Consensus clustering appears to improve the robustness and quality of…
Stephen Coleman, Paul D.W. Kirk, Chris Wallace
Cluster analysis is an integral part of precision medicine and systems biology, used to define groups of patients or biomolecules. However, problems such as choosing the number of clusters and issues with high dimensional data arise consistently. An ensemble approach, such as consensus clustering, can overcome some of…
Shaina Race, Carl D. Meyer
A novel framework for consensus clustering is presented which has the ability to determine both the number of clusters and a final solution using multiple algorithms. A consensus similarity matrix is formed from an ensemble using multiple algorithms and several values for k. A variety of dimension reduction techniques…
Behnam Yousefi, Benno Schwikowski
Clustering plays an important role in a multitude of bioinformatics applications, including protein function prediction, population genetics, and gene expression analysis. The results of most clustering algorithms are sensitive to variations of the input data, the clustering algorithm and its parameters, and individual…
Deguang Kong, Miao Lü, Konstantin Shmakov, Jian Yang
Consensus clustering aggregates partitions in order to find a better fit by reconciling clustering results from different sources/executions. In practice, there exist noise and outliers in clustering task, which, however, may significantly degrade the performance. To address this issue, we propose a novel algorithm –…
Yasin Senbabaoglu, George Michailidis, Jun Z. Li
Consensus clustering (CC) is an unsupervised class discovery method widely used to study sample heterogeneity in high-dimensional datasets. It calculates “consensus rate” between any two samples as how frequently they are grouped together in repeated clustering runs under a certain degree of random perturbation. The…
Faisal Saeed, Naomie Salim, Ammar Abdo
Background Although many consensus clustering methods have been successfully used for combining multiple classifiers in many areas such as machine learning, applied statistics, pattern recognition and bioinformatics, few consensus clustering methods have been applied for combining multiple clusterings of chemical…
Shouvick Mondal, Arko Banerjee
— Recently ensemble selection for consensus clustering has emerged as a research problem in Machine Intelligence. Normally consensus clustering algorithms take into account the entire ensemble of clustering, where there is a tendency of generating a very large size ensemble before computing its consensus. One can avoid…
Luzie Helfmann, Johannes von Lindheim, Mattes Mollenhauer, Ralf Banisch
'Ralf Banisch'] Quality assessments of models in unsupervised learning and clustering verification in particular have been a long-standing problem in the machine learning research. The lack of robust and universally applicable cluster validity scores often makes the algorithm selection and hyperparameter evaluation a…
Mimi Zhang
Clustering ensemble has emerged as a powerful tool for improving both the robustness and the stability of results from individual clustering methods. Weighted clustering ensemble arises naturally from clustering ensemble. One of the arguments for weighted clustering ensemble is that elements (clusterings or clusters)…
Antoine Lacour, Hamza Ibrahim, Andrea Volkamer, Anna K. H. Hirsch
In this study, we introduce DockM8, an innovative open-source platform designed for consensus virtual screening in drug design. Leveraging various docking algorithms and scoring functions, DockM8 provides a highly customizable workflow for structure-based virtual screening. In rigorous evaluations across the DEKOIS…
Edgar López-López, José L. Medina-Franco
Drug-induced liver injury (DILI) is the principal reason for failure in developing drug candidates. It is the most common reason to withdraw from the market after a drug has been approved for clinical use. Therefore, a current challenge is enhancing the accuracy of DILI events' predictive models. In this context, data…
Mary Pitman, David Hahn, Gary Tresadern, David Mobley
Drug discovery is accelerated with computational methods such as alchemical simulations to estimate ligand affinities. In particular, relative binding free energy (RBFE) simulations are beneficial for lead optimization. To use RBFE simulations to compare prospective ligands in silico, researchers first plan the…
Evangelos Karatzas, Maria Gkonta, Joana Hotova, Fotis A. Baltoumas + 4 more
Clustering is the process of grouping together different data objects based on similar properties. Clustering has applications in various case studies from several fields such as graph theory, image analysis, pattern recognition, statistics and others. Nowadays, there are numerous algorithms and tools able to generate…
Lori Dalton, Virginia Ballarin, Marcel Brun
The development of microarray technology has enabled scientists to measure the expression of thousands of genes simultaneously, resulting in a surge of interest in several disciplines throughout biology and medicine. While data clustering has been used for decades in image processing and pattern recognition, in recent…
Hang Hu, Jyothsna Padmakumar Bindu, Julia Laskin
Mass spectrometry imaging (MSI) is widely used for the label-free molecular mapping of biological samples. The identification of co-localized molecules in MSI data is crucial to the understanding of biochemical pathways. However, complex MSI data are too large for manual annotation but too small for training deep…
Authors not listed
This work investigates different formulations of internal Coordinates for molecular dynamics (MD) simulations. The goal is to assess their advantages and limitations. Furthermore, a method is presented that evaluates the quality of the partitioning of molecular structural data into clusters based on statistical…
Authors not listed
Recent advances in artificial intelligence have significantly improved spectral data analysis. In this study, we used unsupervised machine learning to classify chemical compounds based on infrared (IR) spectral images, without relying on prior chemical knowledge. The potential of machine learning for chemical…
Authors not listed
We present a gridless framework for computing high-dimensional conformational free energy surfaces (FES) of flexible molecules using enhanced sampling trajectories. By combining concurrent well-tempered metadynamics with Density Peaks Advanced (DPA) clustering, our approach bypasses the dimensionality limitations of…