26 papers · ranked by Valyu relevance
Rita Kukafka, Kushan Silva, Seyed Mohammad Ayyoubzadeh, Xiaoguang Lyu + 5 more
'Xiaoguang Lyu' 'Adrian Ahne' 'Guy Fagherazzi' 'Xavier Tannier' 'Thomas Czernichow' 'Francisco Orchard'] Background The amount of available textual health data such as scientific and biomedical literature is constantly growing and becoming more and more challenging for health professionals to properly summarize those…
Phillipe R. Sampaio, Helene Maxcici
We study unsupervised clustering of documents at both the category and template levels using frozen multimodal encoders and classical clustering algorithms. We systematize a model-agnostic pipeline that (i) projects heterogeneous last-layer states from text–layout–vision encoders into tokentype–aware document vectors…
Laurence Hirsch, Robin Hirsch, Bayode Ogunleye
Text clustering holds significant value across various domains due to its ability to identify patterns and group related information. Current approaches which rely heavily on a computed similarity measure between documents are often limited in accuracy and interpretability. We present a novel approach to the problem…
Christine A. Luzon, Luisito Lolong Lacatan, Harold Y. Bangalisan, Jayvee M. Osapdin
Method – In this paper a novel approach has been introduced for search results clustering that is based on the semantics of the retrieved documents rather than the syntax of the terms in those documents. Data clustering was used to improve the information retrieval from the collection of documents. Data were processed…
Suganya Selvaraj, Eunmi Choi, Giovanni Betta
Text document clustering is one of the data mining techniques used in many real-world applications such as information retrieval from IoT Sensors data, duplicate content detection, and document organization. Swarm intelligence (SI) algorithms are suitable for solving complex text document clustering problems compared…
Xin Du, Kumiko Tanaka‐Ishii
We present generative clustering (GC) for clustering a set of documents, X, by using texts Y generated by large language models (LLMs) instead of by clustering the original documents X. Because LLMs provide probability distributions, the similarity between two documents can be rigorously defined in an…
Bartłomiej Starosta, Mieczysław A. Kłopotek, Sławomir T. Wierzchoń, Dariusz Czerski + 3 more
'Dariusz Czerski' 'Marcin Sydow' 'Piotr Borkowski' 'Tiago Pereira'] Spectral clustering methods are known for their ability to represent clusters of diverse shapes, densities etc. However, the results of such algorithms, when applied e.g. to text documents, are hard to explain to the user, especially due to embedding…
Luis Felipe Gutiérrez, Neda Tavakoli, Sima Siami-Namini, Akbar Siami Namin
The coronavirus pandemic has already caused plenty of severe problems for humanity and the economy. The exact impact of the COVID-19 pandemic is still unknown, and economists and financial advisers are exploring all possible scenarios to mitigate the risks arising from the pandemic. An intriguing question is whether…
Justin K. Miller, Tristram J. Alexander
Clustering short text is a difficult problem, owing to the low word co-occurrence between short text documents. This work shows that large language models (LLMs) can overcome the limitations of traditional clustering approaches by generating embeddings that capture the semantic nuances of short text. In this study…
Mieczysław A. Kłopotek, Sławomir T. Wierzchoń, Bartłomiej Starosta, Piotr Borkowski + 2 more
In a previous paper, we proposed an introduction to the explainability of Graph Spectral Clustering results for textual documents, given that document similarity is computed as cosine similarity in term vector space. In this paper, we generalize this idea by considering other embeddings of documents, in particular…
Fernando Simeone, Maik Olher Chaves, Ahmed Ali Abdalla Esmin
The growth in Internet usage has contributed to a large volume of continuously available data, and has created the need for automatic and efficient organization of the data. In this context, text clustering techniques are significant because they aim to organize documents according to their characteristics. More…
Stijn van Dongen
Clustering (a large class of methods) is a standard and often used approach in data analysis for separating data into groups, called clusters, often in large-scale high-dimensional data. A clustering (a data structure) is a partitioning of the data into disjoint clusters. This is sometimes called a flat clustering to…
Junlong Liu, Jiaming Xiao, Xunwen Su, Yonglin Wang
Protein clustering and classification are critical for understanding protein functions and interactions, particularly within structure-based predictions. Traditional sequence-based clustering often overlooks the pivotal role of tertiary structure in determining protein function. Structural clustering remains limited…
Behnam Yousefi, Benno Schwikowski
Clustering plays an important role in a multitude of bioinformatics applications, including protein function prediction, population genetics, and gene expression analysis. The results of most clustering algorithms are sensitive to variations of the input data, the clustering algorithm and its parameters, and individual…
Polina Bombina, Dwayne Tally, Zachary B. Abrams, Kevin R. Coombes
Unsupervised clustering is an important task in biomedical science. We developed a new clustering method, called SillyPutty, for unsupervised clustering. As test data, we generated a series of datasets using the Umpire R package. Using these datasets, we compared SillyPutty to several existing algorithms using multiple…
Ali Turfah, Xiaoquan Wen
Cluster analysis is a widely used unsupervised learning technique in genomic data analysis, with critical applications such as inferring genetic population structures and annotating cell types from single-cell RNA-seq data. However, most existing clustering methods focus on identifying a single optimal partition while…
Parichit Sharma, Sarthak Mishra, Hasan Kurban, Mehmet Dalkilic
This paper introduces p-ClustVal, a novel data transformation technique inspired by p-adic number theory that significantly enhances cluster discernibility in genomics data, specifically Single Cell RNA Sequencing (scRNASeq). By leveraging p-adic-valuation, p-ClustVal integrates with and augments widely used clustering…
Yijia Li, Jonathan Nguyen, David Anastasiu, Edgar A. Arriaga
With the aim of analyzing large-sized multidimensional single-cell datasets, we are describing our method for Cosine-based Tanimoto similarity-refined graph for community detection using Leiden’s algorithm (CosTaL). As a graph-based clustering method, CosTaL transforms the cells with high-dimensional features into a…
Shiya Zhou, Muhammad Asif
Water resource accounting constitutes a fundamental approach for implementing sophisticated management of basin water resources. The quality of water plays a pivotal role in determining the liabilities associated with these resources. Evaluating the quality of water facilitates the computation of water resource…
Lionel Zoubritzky, François-Xavier Coudert
We present here an open-source Julia library for the topological identification of crystalline materials, with algorithmic and computational improvements over the previously available software in the field, resulting in a speed increase of one order of magnitude. This new algorithm and implementation can therefore be…
Qinghua Wang, Jonathan Olshin, K. Vijay-Shanker, Cathy Wu
Chinese hamster ovary (CHO) cells are widely used for mass production of therapeutic proteins in the pharmaceutical industry. With the growing need in optimizing the performance of producer CHO cell lines, research on CHO cell line development and bioprocess continues to increase in recent decades. Bibliographic…
Authors not listed
The screening of chemical libraries is an essential starting point in the drug discovery process. While some researchers desire a more thorough screening of drug targets against a narrower scope of molecules, it is not uncommon for diverse screening sets to be favored during early stages of drug discovery. However, a…
Authors not listed
Recent advances in artificial intelligence have significantly improved spectral data analysis. In this study, we used unsupervised machine learning to classify chemical compounds based on infrared (IR) spectral images, without relying on prior chemical knowledge. The potential of machine learning for chemical…
Authors not listed
This work investigates different formulations of internal Coordinates for molecular dynamics (MD) simulations. The goal is to assess their advantages and limitations. Furthermore, a method is presented that evaluates the quality of the partitioning of molecular structural data into clusters based on statistical…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European and US chemical agencies have listed approximately 800k chemicals where knowledge on potential risks to human health and the environment are lacking. Filling these data gaps experimentally is impossible so in-silico approaches and prediction are essential. Many existing models are however limited by…
Pulan Yu
Associative classification mining (ACM) integrating association rule mining and classification has become a significant tool for knowledge discovery, especially in the chemical domain. Its major advantage is providing high accuracy as well as chemically interpretable models. Additionally, it is able to find…