22 papers · ranked by Valyu relevance
Prabhakar Chalise, Brooke L. Fridley, Shyamal D Peddada
Integrative analyses of high-throughput ‘omic data, such as DNA methylation, DNA copy number alteration, mRNA and protein expression levels, have created unprecedented opportunities to understand the molecular basis of human disease. In particular, integrative analyses have been the cornerstone in the study of cancer…
Evelina Gabasova, John Reid, Lorenz Wernisch, Quaid Morris
Integrative clustering is used to identify groups of samples by jointly analysing multiple datasets describing the same set of biological samples, such as gene expression, copy number, methylation etc. Most existing algorithms for integrative clustering assume that there is a shared consistent set of clusters across…
Evelina Gabasova, John Reid, Lorenz Wernisch
Integrative clustering is used to identify groups of samples by jointly analysing multiple datasets describing the same set of biological samples, such as gene expression, copy number, methylation etc. Most existing algorithms for integrative clustering assume that there is a shared consistent set of clusters across…
Kristoffer H. Hellton, Magne Thoresen
When measuring a range of different genomic, epigenomic, transcriptomic and other variables, an integrative approach to analysis can strengthen inference and give new insights. This is also the case when clustering patient samples, and several integrative cluster procedures have been proposed. Common for these…
Alessandra Cabassi, Paul D W Kirk, Jinbo Xu
Thanks to technological advances, both the availability and the diversity of omic datasets have hugely increased in recent years (). These datasets provide information on multiple levels of biological systems, going from the genomic and epigenomic level, to gene and protein expression level, up to the metabolomic…
Benjamin Balluff, Achim Buck, Marta Martin‐Lorenzo, Frédéric Dewez + 4 more
Clustering algorithms are powerful data analysis methods to reveal the inherent relation of objects in a high-dimensional feature space. Large, high-dimensional datasets are now characteristic of most biomedical research owing to the widespread use of next-generation sequencing and high-throughput mass spectrometry.…
Ronglai Shen, Sijian Wang, Qianxing Mo
High resolution microarrays and second-generation sequencing platforms are powerful tools to investigate genome-wide alterations in DNA copy number, methylation and gene expression associated with a disease. An integrated genomic profiling approach measures multiple omics data types simultaneously in the same set of…
Alessandra Cabassi, Paul Kirk
Summary: Diverse applications – particularly in tumour subtyping – have demonstrated the importance of integrative clustering techniques for combining information from multiple data sources. Cluster Of Clusters Analysis (COCA) is one such approach that has been widely applied in the context of tumour subtyping.…
Bastian Pfeifer, Michael G. Schimek
Recent advances in multi-omics clustering methods enable a more fine-tuned separation of cancer patients into clinical relevant clusters. These advancements have the potential to provide a deeper understanding of cancer progression and may facilitate the treatment of cancer patients. Here, we present a simple…
David M. Swanson, Tonje Lien, Helga Bergholtz, Therese Sørlie + 1 more
Unsupervised clustering is important in disease subtyping, among having other genomic applications. As genomic data has become more multifaceted, how to cluster across data sources for more precise subtyping is an ever more important area of research. Many of the methods proposed so far, including iCluster and Cluster…
Alessandra Cabassi, Sylvia Richardson, Paul Kirk
Summary: When using Markov chain Monte Carlo (MCMC) algorithms to perform inference for Bayesian clustering models, such as mixture models, the output is typically a sample of clusterings (partitions) drawn from the posterior distribution. In practice, a key challenge is how to summarise this output. Here we build upon…
Galadriel Brière, Élodie Darbo, Patricia Thébault, Raluca Uricaru
Facing the diversity of omic data and the difficulty of selecting one result over all those produced by several methods, consensus strategies have the potential to reconcile multiple inputs and to produce robust results. Here, we introduce ClustOmics, a generic consensus clustering tool that we use in the context of…
Minjie Wang, Genevera I. Allen
In mixed multi-view data, multiple sets of diverse features are measured on the same set of samples. By integrating all available data sources, we seek to discover common group structure among the samples that may be hidden in individualistic cluster analyses of a single data-view. While several techniques for such…
Maryam Pouryahya, Jung Hun Oh, Pedram Javanmard, James C. Mathews + 3 more
The remarkable growth of multi-platform genomic profiles has led to the multiomics data integration challenge. The effective integration of such data provides a comprehensive view of the molecular complexity of cancer tumors and can significantly improve clinical out-come predictions. In this study, we present a novel…
Anna Hristoskova, Veselka Boeva, Elena Tsiporkova
Background Presently, with the increasing number and complexity of available gene expression datasets, the combination of data from multiple microarray studies addressing a similar biological question is gaining importance. The analysis and integration of multiple datasets are expected to yield more reliable and robust…
Matthías Kormáksson, James G. Booth, María E. Figueroa, Ari Melnick
> In many fields, researchers are interested in large and complex biological processes. Two important examples are gene expression and DNA methylation in genetics. One key problem is to identify aberrant patterns of these processes and discover biologically distinct groups. In this article we develop a model-based…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Himaghna Bhattacharjee, Jackson Burns, Dionisios Vlachos
The recent advances in deep learning, generative modeling, and statistical learning have ushered in a renewed interest in traditional cheminformatics tools and methods. Quantifying molecular similarity is essential in molecular generative modeling, exploratory molecular synthesis campaigns, and drug-discovery…
Hang Hu, Jyothsna Padmakumar Bindu, Julia Laskin
Mass spectrometry imaging (MSI) is widely used for the label-free molecular mapping of biological samples. The identification of co-localized molecules in MSI data is crucial to the understanding of biochemical pathways. However, complex MSI data are too large for manual annotation but too small for training deep…
Authors not listed
Recent advances in artificial intelligence have significantly improved spectral data analysis. In this study, we used unsupervised machine learning to classify chemical compounds based on infrared (IR) spectral images, without relying on prior chemical knowledge. The potential of machine learning for chemical…
Kamen Petrov, Andreas Bender
Organizing and partitioning sets of chemical structures is of considerable practical significance e.g. in compound library analysis and the post-processing of screening hit lists. Approaches such as unsupervised clustering are computationally demanding and dataset-dependent; on the other hand, rule-based methods, such…
Authors not listed
Mass spectrometry (MS) is a cornerstone technology in modern molecular biology, powering diverse applications across proteomics, metabolomics, lipidomics, glycomics, and beyond. As the field continues to evolve, rapid advancements in instrumentation, acquisition strategies, machine learning, and scalable computing have…