Search · four archives
Search · four archives
24 papers · ranked by Valyu relevance
Di Li, Qinglin Mei, Guojun Li
Single-cell RNA sequencing (scRNA-seq) technologies have been driving the development of algorithms of clustering heterogeneous cells. We introduce a novel clustering algorithm scQA, which can effectively and efficiently recognize different cell types via qualitative and quantitative analysis. It iteratively extracts…
Diego A. Camacho-Hernández, Victor E. Nieto-Caballero, José E. León-Burguete, Julio A. Freyre-González
Identifying groups that share common features among datasets through clustering analysis is a typical problem in many fields of science, particularly in post-omics and systems biology research. In respect of this, quantifying how a measure can cluster or organize intrinsic groups is important since currently there is…
Onofrio Rosario Battaglia, Benedetto Di Paola, Claudio Fazio
In the last years many studies examined the consistency of students' answers in a variety of contexts. Some of these papers tried to develop more detailed models of the consistency of students' reasoning, or to subdivide a sample of students into intellectually similar subgroups. The problem of taking a set of data and…
André P. Neto-Bradley, Rishika Rangarajan, Ruchi Choudhary, Amir B. Bazaz
'Amir B. Bazaz'] Studies on clean energy transition amongst low-income urban households in the Global South use an array of qualitative and quantitative methods. However, attempts to combine qualitative and quantitative methods are rare and there are a lack of systematic approaches to this. This paper demonstrates a…
Marsha A Wilcox, Diego F Wyszynski, Carolien I Panhuysen, Qianli Ma + 3 more
'Agustin Yip' 'John Farrell' 'Lindsay A Farrer'] Background The Framingham Heart Study has contributed a great deal to advances in medicine. Most of the phenotypes investigated have been univariate traits (quantitative or qualitative). The aims of this study are to derive multivariate traits by identifying homogeneous…
Diego A. Camacho-Hernández, Victor E. Nieto-Caballero, José E. León-Burguete, Julio A. Freyre-González
'José E. León-Burguete' 'Julio A. Freyre-González'] Abstract: Identifying groups that share common features among datasets through clustering analysis is a typical problem in many fields of science, particularly in post-omics and systems biology research. In respect of this, quantifying how a measure can cluster or…
Yun‐Cheng Tsai, Yen-Ku Liu, Samuel Yen-Chi Chen
Blockchain transaction data is inherently high-dimensional, noisy, and entangled, posing substantial challenges for traditional clustering algorithms. While quantum-enhanced clustering models have demonstrated promising performance gains, their interpretability remains limited, restricting their application in…
Chenyue W. Hu, Steven M. Kornblau, John H. Slater, Amina A. Qutub
Estimating the optimal number of clusters is a major challenge in applying cluster analysis to any type of dataset, especially to biomedical datasets, which are high-dimensional and complex. Here, we introduce an improved method, Progeny Clustering, which is stability-based and exceptionally efficient in computing, to…
Christian Hennig, Cinzia Viroli, Laura Anderlucci
A new cluster analysis method, K-quantiles clustering, is introduced. K-quantiles clustering can be computed by a simple greedy algorithm in the style of the classical Lloyd's algorithm for K-means. It can be applied to large and high-dimensional datasets. It allows for within-cluster skewness and internal variable…
Linda Vidman, David Källberg, Patrik Rydén
Clustering of gene expression data is widely used to identify novel subtypes of cancer. Plenty of clustering approaches have been proposed, but there is a lack of knowledge regarding their relative merits and how data characteristics influence the performance. We evaluate how cluster analysis choices affect the…
J W G Addy, J Langhorne
Assessing how adequate clusters fit a dataset and finding an optimum number of clusters is a difficult process. A membership matrix and the degree of membership matrix is suggested to determine the homogeneity of a cluster fit. Maximisation of the ratio of the overall degree of membership at cluster number lag 1 is…
Misa Goudo, Masahiro Sugimoto, Satoru Hiwa, Tomoyuki Hiroyasu
In processing metabolomics data, multidimensional quantitative data from thousands of metabolites are often sparse, that is, only a small fraction of metabolites are relevant to the phenotype of interest. Clustering is therefore used to discover subtypes from omics data. Sparse processing, which selects important…
So Hyeon Bak, Hye Yun Park, Jin Hyun Nam, Ho Yun Lee + 4 more
In the present study, quantitative CT features dealing both with fibrotic score and emphysema index were used for clustering of radiologic phenotyping in patients with IPF, yielding three clusters. The radiologic phenotypic subgroups identified using cluster analysis according quantitative CT features differed…
Michael Kern, Alexander Lex, Nils Gehlenborg, Chris R. Johnson
With ever-increasing amounts of data produced in biology research, scientists are in need of efficient data analysis methods. Cluster analysis, combined with visualization of the results, is one such method that can be used to make sense of large data volumes. At the same time, cluster analysis is known to be imperfect…
Diogo Seca, João Mendes‐Moreira, Tiago Mendes-Neves, Ricardo Brandão da Silva Sousa
'Ricardo Brandão da Silva Sousa'] Clustering can be used to extract insights from data or to verify some of the assumptions held by the domain experts, namely data segmentation. In the literature, few methods can be applied in clustering qualitative values using the context associated with other variables present in…
Simon Crase, Suresh N. Thennadil, Usman Qamar
Cluster analysis is a valuable unsupervised machine learning technique that is applied in a multitude of domains to identify similarities or clusters in unlabelled data. However, its performance is dependent of the characteristics of the data it is being applied to. There is no universally best clustering algorithm…
Cyril Esnault, Melissa Rollot, Pauline Guilmin, Jean-Daniel Zucker
The exploration of heath data by clustering algorithms allows to better describe the populations of interest by seeking the sub-profiles that compose it. This therefore reinforces medical knowledge, whether it is about a disease or a targeted population in real life. Nevertheless, contrary to the so-called conventional…
Sudhakar Jonnalagadda, Rajagopalan Srinivasan
Background Clustering techniques are routinely used in gene expression data analysis to organize the massive data. Clustering techniques arrange a large number of genes or assays into a few clusters while maximizing the intra-cluster similarity and inter-cluster separation. While clustering of genes facilitates…
Authors not listed
Recent advances in artificial intelligence have significantly improved spectral data analysis. In this study, we used unsupervised machine learning to classify chemical compounds based on infrared (IR) spectral images, without relying on prior chemical knowledge. The potential of machine learning for chemical…
Jakub Kubečka, Vitus Besel, Ivo Neefjes, Yosef Knattrup + 3 more
Computational modeling of atmospheric molecular clusters requires a comprehensive understanding of their complex configurational spaces, interaction patterns, stabilities against fragmentation, and even dynamic behaviors. To address these needs, we introduce the Jammy Key framework, a collection of automated scripts…
Authors not listed
The screening of chemical libraries is an essential starting point in the drug discovery process. While some researchers desire a more thorough screening of drug targets against a narrower scope of molecules, it is not uncommon for diverse screening sets to be favored during early stages of drug discovery. However, a…
Mingze Bai, Jingwen Deng, Chengxin Dai, Julianus Pfeuffer + 1 more
Testing for significant differences in quantities on protein level is a common goal of many LFQ-based mass spectrometry proteomics experiments. Starting from a table of protein and/or peptide quantities from a fixed proteomics quantification software, there exists a multitude of tools and R packages to perform the…
Bartłomiej Fliszkiewicz, Marcin Sajdak
The aim of the following research is to assess the applicability of calculated quantum properties of molecular fragments as molecular descriptors in machine learning classification task. The research is based on bio-concentration and QM9-extended databases. A number of compounds with results from quantum-chemical…
Elizaveta I. Shestoperova, Daniil G. Ivanov, Eric R. Strieter
The diversity of ubiquitin modifications calls for methods to better characterize ubiquitin chain linkage, length, and morphology. Here, we use multiple linear regression analysis coupled with ion mobility mass spectrometry (IM-MS) to quantify the relative abundance of different ubiquitin dimer isomers. We demonstrate…