Search · four archives
Search · four archives
27 papers · ranked by Valyu relevance
Yuansong Zeng, Zhuoyi Wei, Fengqi Zhong, Zixiang Pan + 2 more
Clustering analysis is widely utilized in single-cell RNA-sequencing (scRNA-seq) data to discover cell heterogeneity and cell states. While many clustering methods have been developed for scRNA-seq analysis, most of these methods require to provide the number of clusters. However, it is not easy to know the exact…
Claudia Plant, Lena G. M. Bauer, Christian Böhm
How to find a natural grouping of a large real data set? Clustering requires a balance between abstraction and representation. To identify clusters, we need to abstract from superfluous details of individual objects. But we also need a rich representation that emphasizes the key features shared by groups of objects…
Georgios Vardakas, Ioannis Papakostas, Aristidis Likas
Well-Separated Clusters Authors: ['Georgios Vardakas' 'Ioannis Papakostas' 'Aristidis Likas'] Unsupervised learning has gained prominence in the big data era, offering a means to extract valuable insights from unlabeled datasets. Deep clustering has emerged as an important unsupervised category, aiming to exploit the…
Gang Wu, Junjun Jiang, Xianming Liu
Single-cell RNA sequencing (scRNA-seq) reveals the heterogeneity and diversity among individual cells and allows researchers conduct cell-wise analysis. Clustering analysis is a fundamental step in analyzing scRNA-seq data which is needed in many downstream tasks. Recently, some deep clustering based methods exhibit…
Kart–Leong Lim
—Deep clustering is a recent deep learning technique which combines deep learning with traditional unsupervised clustering. At the heart of deep clustering is a loss function which penalizes samples for being an outlier from their ground truth cluster centers in the latent space. The probabilistic variant of deep…
Ziyou Zheng, Shuzhen Zhang, Hailong Song, Qi Yan
Deep clustering has been widely applicated in various fields, including natural image and language processing. However, when it is applied to hyperspectral image (HSI) processing, it encounters challenges due to high dimensionality of HSI and complex spatial-spectral characteristics. This study introduces a kind of…
Charles A. Ellis, Robyn L. Miller, Vince D. Calhoun
Machine learning methods have frequently been applied to electroencephalography (EEG) data. However, while supervised EEG classification is well-developed, relatively few studies have clustered EEG, which is problematic given the potential for clustering EEG to identify novel subtypes or patterns of dynamics that could…
Eugen-Richard Ardelean, Raluca Laura Portase
Spike sorting is the process of identifying the source neurons for neuronal activity recorded from extracellular electrodes. Traditional spike sorting pipelines separate the process into distinct feature extraction and clustering steps, which may not optimally capture the complex structure of spike data. This study…
Ahmed Salah, David Yevick
– This paper introduces a modified variational autoencoder (VAEs) that contains an additional neural network branch. The resulting "branched VAE" (BVAE) contributes a classification component based on the class labelsto the total loss and therefore imparts categorical information to the latent representation. As a…
Yunhe Wang, Zhuohan Yu, Shaochuan Li, Chuang Bian + 4 more
The cell is the basic unit of growth and development of an organism and has unique biological functions. The heterogeneity between cells in a cell population has isogenic properties, which can ascend from stochastic expression of genes, proteins and metabolites (). Conventional bulk RNA sequencing (RNA-seq) averages…
Jun Seo Ha, Hyundoo Jeong, Achraf El Allali
Recent advances in single-cell sequencing techniques have enabled gene expression profiling of individual cells in tissue samples so that it can accelerate biomedical research to develop novel therapeutic methods and effective drugs for complex disease. The typical first step in the downstream analysis pipeline is…
Miaozhuang Cai, Yin Zheng, Zhengyang Peng, Chunyan Huang + 2 more
'Haoxia Jiang' 'Luan Carlos de Sena Monteiro Ozelim'] Time series data complexity presents new challenges in clustering analysis across fields such as electricity, energy, industry, and finance. Despite advances in representation learning and clustering with Variational Autoencoders (VAE) based deep learning…
Debapriya Roy
Clustering is a long-standing problem area in data mining. The centroid-based classical approaches to clustering mainly face difficulty in the case of high dimensional inputs such as images. With the advent of deep neural networks, a common approach to this problem is to map the data to some latent space of…
Mona Suliman AlZuhair, Mohamed Maher Ben Ismail, Ouiem Bchir, Adrian Barbu
'Adrian Barbu'] Semi-supervised clustering can be viewed as a clustering paradigm that exploits both labeled and unlabeled data to steer learning accurate data clusters and avoid local minimum solutions. Nonetheless, the attempts to refine existing semi-supervised clustering methods are relatively limited when compared…
Kart–Leong Lim
Jensen-Shannon Divergence Clustering Loss Authors: ['Kart–Leong Lim'] Abstract—Deep clustering is an emerging topic in deep learning where traditional clustering is performed in deep learning feature space. However, clustering and deep learning are often mutually exclusive. In the autoencoder based deep clustering, the…
Teja Potu, Yunfei Hu, Rituparna Khan, Srinija Dharani + 4 more
Intra-tumor heterogeneity (ITH) is a compounding factor for cancer prognosis and treatment. Single-cell DNA sequencing (scDNA-seq) provides cellular resolution of the variations in a cell and has been widely used to study cancer progression and responses to drug and treatment. While the low coverage scDNA-seq…
Zhanwen Cheng, Feijiang Li, Jieting Wang, Yuhua Qian
Deep clustering methods improve the performance of clustering tasks by jointly optimizing deep representation learning and clustering. While numerous deep clustering algorithms have been proposed, most of them rely on artificially constructed pseudo targets for performing clustering. This construction process requires…
Guanfang Dong, Chenqiu Zhao, Anup Basu
—Distribution learning focuses on learning the probability density function from a set of data samples. In contrast, clustering aims to group similar objects together in an unsupervised manner. Usually, these two tasks are considered unrelated. However, the relationship between the two may be indirectly correlated…
Sangyeon Lee, Hanjin Kim, Doheon Lee
Regression analysis is one of the most widely applied methods in many fields including bio-medical study. Dimensionality reduction is also widely used for data preprocessing and feature selection analysis, to extract high-impact features from the predictions. As the complexity of both data and prediction models…
Adrian Wheeldon, Alexander Serb
Latent representations are a necessary component of cognitive artificial intelligence (AI) systems. Here, we investigate the performance of various sequential clustering algorithms on latent representations generated by autoencoder and convolutional neural network (CNN) models. We also introduce a new algorithm, called…
Xiang Lin, Jianlan Ren, Le Gao, Zhi Wei + 1 more
The scRNA-seq technology enables high-resolution profiling and analysis of individual cells. The increasing availability of datasets and advancements in technology have prompted researchers to integrate existing annotated datasets with newly sequenced datasets for a more comprehensive analysis. It is important to…
Authors not listed
Acoustic measurements of batteries are known to be correlated to their state-of-charge, creating opportunities for state estimation that do not rely on electrical signals. State estimators are typically parametric models fitted from data, often from the broad toolbox of machine learning. Such models can be easily…
Samuel Renaud, Rachael Mansbach
Current antibacterial treatments cannot overcome the rapidly growing resistance of bacteria to antibiotic drugs, and novel treatment methods are required. One option is the development of new antimicrobial peptides (AMPs), to which bacterial resistance build-up is comparatively slow. Deep generative models have…
Tagir Akhmetshin, Arkadii Lin, Timur Madzhidov, Alexandre Varnek
Autoencoders represent a promising technique for the inverse quantitative structure-activity relationship (QSAR) task. However, undesirable bias, such as atom ordering, affects the neighbourhood behaviour of autoencoders’ latent space and, consequently, usage of the latent vectors as variables in machine-learning…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Pavel Kohout, Michal Vasina, Marika Majerova, Veronika Novakova + 5 more
Enzymes play a crucial role in sustainable industrial applications, with their optimization posing a formidable challenge due to the intricate interplay among residues. Computational methodologies predominantly rely on evolutionary insights, leveraging homologous sequences to pinpoint conserved and functionally…
Authors not listed
Compound similarity is fundamental to various cheminformatics analyses, particularly in the drug discovery industry, where the structure-activity principle is central to medicinal chemistry. Historically, binary fingerprints combined with Tanimoto and “Tanimoto-related metrics” (such as Dice, Sørensen–Dice, and…