25 papers · ranked by Valyu relevance
Christopher M. De Vries, Shlomo Geva, Andrew Trotman
Divergence from a random baseline is a technique for the evaluation of document clustering. It ensures cluster quality measures are performing work that prevents ineffective clusterings from giving high scores to clusterings that provide no useful result. These concepts are defined and analysed using intrinsic and…
Kevin W. Boyack, David Newman, Russell J. Duhon, Richard Klavans + 7 more
'Michael Patek' 'Joseph R. Biberstine' 'Bob Schijvenaars' 'André Skupin' 'Nianli Ma' 'Katy Börner' 'Colin Allen'] Background We investigate the accuracy of different similarity approaches for clustering over two million biomedical documents. Clustering large sets of text documents is important for a variety of…
Jens Dörpinghaus, Sebastian Schaaf, Marc Jacobs
Document clustering is widely used in science for data retrieval and organisation. DocClustering is developed to include and use a novel algorithm called PS-Document Clustering that has been first introduced in 2017. This method combines approaches of graph theory with state of the art NLP-technologies. This new…
Muhammad Rafi, Farnaz Amin, M. Shahid Shaikh
Document clustering is an unsupervised approach in which a large collection of documents (corpus) is subdivided into smaller, meaningful, identifiable, and verifiable sub-groups (clusters). Meaningful representation of documents and implicitly identifying the patterns, on which this separation is performed, is the…
Jens Dörpinghaus, Sebastian Schaaf, Marc Jacobs
Background In text mining, document clustering describes the efforts to assign unstructured documents to clusters, which in turn usually refer to topics. Clustering is widely used in science for data retrieval and organisation. Results In this paper we present and discuss a novel graph-theoretical approach for document…
R. Jensi, Wiselin Jiji G
Text Document Clustering is one of the fastest growing research areas because of availability of huge amount of information in an electronic form. There are several number of techniques launched for clustering documents in such a way that documents within a cluster have high intra-similarity and low inter-similarity to…
Suganya Selvaraj, Eunmi Choi, Gianvito Pio, Roberto Corizzo + 1 more
'Michelangelo Ceci'] Text document clustering refers to the unsupervised classification of textual documents into clusters based on content similarity and can be applied in applications such as search optimization and extracting hidden information from data generated by IoT sensors. Swarm intelligence (SI) algorithms…
Amjad F. Alsuhaim, Aqil M. Azmi, Muhammad Hussain, Friedhelm Schwenker
'Friedhelm Schwenker'] Traditional information retrieval systems return a ranked list of results to a user’s query. This list is often long, and the user cannot explore all the results retrieved. It is also ineffective for a highly ambiguous language such as Arabic. The modern writing style of Arabic excludes the…
Khaled Abdalgader, Atheer A. Matroud, Khaled Hossin, Xiangjie Kong
Sentence clustering plays a central role in various text-processing activities and has received extensive attention for measuring semantic similarity between compared sentences. However, relatively little focus has been placed on evaluating clustering performance using available similarity measures that adopt…
Xin Du, Kumiko Tanaka‐Ishii
We present generative clustering (GC) for clustering a set of documents, X, by using texts Y generated by large language models (LLMs) instead of by clustering the original documents X. Because LLMs provide probability distributions, the similarity between two documents can be rigorously defined in an…
Vivek Mehta, Seema Bawa, Jasmeet Singh
A massive amount of textual data now exists in digital repositories in the form of research articles, news articles, reviews, Wikipedia articles, and books, etc. Text clustering is a fundamental data mining technique to perform categorization, topic extraction, and information retrieval. Textual datasets, especially…
Rajendra Kumar Roul, Shubham Rohan Asthana, Sanjay K. Sahay
—With the rising quantity of textual data available in electronic format, the need to organize it become a highly challenging task. In the present paper, we explore a document organization framework that exploits an intelligent hierarchical clustering algorithm to generate an index over a set of documents. The…
Christopher M. De Vries, Lance De Vine, Shlomo Geva, Richi Nayak
The proliferation of the web presents an unsolved problem of automatically analyzing billions of pages of natural language. We introduce a scalable algorithm that clusters hundreds of millions of web pages into hundreds of thousands of clusters. It does this on a single mid-range machine using efficient algorithms and…
Hossam M. J. Mustafa, Masri Ayob, Dheeb Albashish, Sawsan Abu-Taleb + 1 more
'Mohd Nadhir Ab Wahab'] The text clustering is considered as one of the most effective text document analysis methods, which is applied to cluster documents as a consequence of the expanded big data and online information. Based on the review of the related work of the text clustering algorithms, these algorithms…
Christine A. Luzon, Luisito Lolong Lacatan, Harold Y. Bangalisan, Jayvee M. Osapdin
Method – In this paper a novel approach has been introduced for search results clustering that is based on the semantics of the retrieved documents rather than the syntax of the terms in those documents. Data clustering was used to improve the information retrieval from the collection of documents. Data were processed…
Sawsan Kanj, Thomas Brüls, Stéphane Gazut
We present a new algorithm to cluster high dimensional sequence data, and its application to the field of metagenomics, which aims to reconstruct individual genomes from a mixture of genomes sampled from an environ-mental site, without any prior knowledge of reference data (genomes) or the shape of clusters. Such…
Evangelos Karatzas, Maria Gkonta, Joana Hotova, Fotis A. Baltoumas + 4 more
Clustering is the process of grouping together different data objects based on similar properties. Clustering has applications in various case studies from several fields such as graph theory, image analysis, pattern recognition, statistics and others. Nowadays, there are numerous algorithms and tools able to generate…
Joao C. Marques, Michael B. Orger
How to partition a data set into a set of distinct clusters is a ubiquitous and challenging problem. The fact that data sets vary widely in features such as cluster shape, cluster number, density distribution, background noise, outliers and degree of overlap, makes it difficult to find a single algorithm that can be…
Polina Bombina, Dwayne Tally, Zachary B. Abrams, Kevin R. Coombes
Unsupervised clustering is an important task in biomedical science. We developed a new clustering method, called SillyPutty, for unsupervised clustering. As test data, we generated a series of datasets using the Umpire R package. Using these datasets, we compared SillyPutty to several existing algorithms using multiple…
Kushal K Dey, Chiaowen Joyce Hsiao, Matthew Stephens
Grade of membership models, also known as “admixture models”, “topic models” or “Latent Dirichlet Allocation”, are a generalization of cluster models that allow each sample to have membership in multiple clusters. These models are widely used in population genetics to model admixed individuals who have ancestry from…
Hang Hu, Jyothsna Padmakumar Bindu, Julia Laskin
Mass spectrometry imaging (MSI) is widely used for the label-free molecular mapping of biological samples. The identification of co-localized molecules in MSI data is crucial to the understanding of biochemical pathways. However, complex MSI data are too large for manual annotation but too small for training deep…
Authors not listed
The screening of chemical libraries is an essential starting point in the drug discovery process. While some researchers desire a more thorough screening of drug targets against a narrower scope of molecules, it is not uncommon for diverse screening sets to be favored during early stages of drug discovery. However, a…
Authors not listed
Recent advances in artificial intelligence have significantly improved spectral data analysis. In this study, we used unsupervised machine learning to classify chemical compounds based on infrared (IR) spectral images, without relying on prior chemical knowledge. The potential of machine learning for chemical…
Authors not listed
This work investigates different formulations of internal Coordinates for molecular dynamics (MD) simulations. The goal is to assess their advantages and limitations. Furthermore, a method is presented that evaluates the quality of the partitioning of molecular structural data into clusters based on statistical…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European and US chemical agencies have listed approximately 800k chemicals where knowledge on potential risks to human health and the environment are lacking. Filling these data gaps experimentally is impossible so in-silico approaches and prediction are essential. Many existing models are however limited by…