22 papers · ranked by Valyu relevance
Matthias Rüdiger, David Antons, Amol M. Joshi, Torsten-Oliver Salge + 1 more
Topic modeling is a popular technique for exploring large document collections. It has proven useful for this task, but its application poses a number of challenges. First, the comparison of available algorithms is anything but simple, as researchers use many different datasets and criteria for their evaluation. A…
Etana Fikadu Dinsa, Mrinal Das, Teklu Urgessa Abebe
Afaan Oromo is a resource-scarce language with limited tools developed for its processing, posing significant challenges for natural language tasks. The tools designed for English do not work efficiently for Afaan Oromo due to the linguistic differences and lack of well-structured resources. To address this challenge…
Eric Austin, Shraddha Makwana, Amine Trabelsi, Christine Largeron + 1 more
Topic modeling aims to discover latent themes in collections of text documents. It has various applications across fields such as sociology, opinion analysis, and media studies. In such areas, it is essential to have easily interpretable, diverse, and coherent topics. An efficient topic modeling technique should…
Hamza H.M. Altarturi, Muntadher Saadoon, Nor Badrul Anuar, Daniel de Oliveira
An immense volume of digital documents exists online and offline with content that can offer useful information and insights. Utilizing topic modeling enhances the analysis and understanding of digital documents. Topic modeling discovers latent semantic structures or topics within a set of digital textual documents.…
Giorgia Minello, Carlo Romano Marcello Alessandro Santagiustina, Massimo Warglien, Fu Lee Wang
During the COVID-19 pandemic, the scientific literature related to SARS-COV-2 has been growing dramatically. These literary items encompass a varied set of topics, ranging from vaccination to protective equipment efficacy as well as lockdown policy evaluations. As a result, the development of automatic methods that…
Hyojin Park, Joachim Gross
Neural representation of lexico-semantics in speech processing has been revealed in recent years. However, to date, how the brain makes sense of the higher-level semantic gist (topic keywords) of a continuous speech remains mysterious. Capitalizing on a generative probabilistic topic modelling algorithm on speech…
Xiaobao Wu, T. Q. Nguyen, Anh Tuan Luu
Topic models have been prevalent for decades to discover latent topics and infer topic proportions of documents in an unsupervised fashion. They have been widely used in various applications like text analysis and context recommendation. Recently, the rise of neural networks has facilitated the emergence of a new…
Dominic B. Dayta, Erniel B. Barrios
Legacy procedures for topic modelling have generally suffered problems of overfitting and a weakness towards reconstructing sparse topic structures. This paper proposes SemiparTM, a two-step approach utilizing nonnegative matrix factorization and semiparametric regression in topic modeling. SemiparTM enables the…
Sergei Koltcov, Anton Surkov, Vladimir Filippov, Vera Ignatenko + 1 more
'Davide Chicco'] Topic modeling is a widely used instrument for the analysis of large text collections. In the last few years, neural topic models and models with word embeddings have been proposed to increase the quality of topic solutions. However, these models were not extensively tested in terms of stability and…
Rebecca Danning, Zheng Tracy Ke, Rong Ma, Xihong Lin
Count data are ubiquitous across many applications in which understanding hidden patterns, or latent structure, is of interest. Topic modeling is a powerful tool for detecting latent structure in count data. However, standard topic modeling methods are often constrained by their restrictive assumptions, susceptible to…
Johannes Schneider
Pre-trained language models have led to a new state-of-the-art in many NLP tasks. However, for topic modeling, statistical generative models such as LDA are still prevalent, which do not easily allow incorporating contextual word vectors. They might yield topics that do not align very well with human judgment. In this…
Uttam Chauhan, Shrusti Shah, Dharati Shiroya, Dipti Solanki + 9 more
'Zeel Patel' 'Jitendra Bhatia' 'Sudeep Tanwar' 'Ravi Sharma' 'Verdes Marina' 'Maria Simona Raboaca' 'Chang Choi' 'Kiho Lim' 'Gyuho Choi'] Topic modeling is a machine learning algorithm based on statistics that follows unsupervised machine learning techniques for mapping a high-dimensional corpus to a low-dimensional…
Alex Gorbulev, Vasiliy Alekseev, Konstantin Vorontsov
Topic modelling is fundamentally a soft clustering problem (of known objects—documents, over unknown clusters—topics). That is, the task is incorrectly posed. In particular, the topic models are unstable and incomplete. All this leads to the fact that the process of finding a good topic model (repeated hyperparameter…
Márton Kardos, Jan Kostkan, Arnault‐Quentin Vermillet, Kristoffer L. Nielbo + 1 more
'Kristoffer L. Nielbo' 'Roberta Rocca'] Topic models are useful tools for discovering latent semantic structures in large textual corpora. Topic modeling historically relied on bag-of-words representations of language. This approach makes models sensitive to the presence of stop words and noise, and does not utilize…
Satyajeet Sahoo, Jhareswar Maiti, Virendra Kumar Tewari
—An important aspect of text mining involves information retrieval in form of discovery of semantic themes (topics) from documents using topic modelling. While generative topic models like Latent Dirichlet Allocation (LDA) elegantly model topics as probability distributions and are useful in identifying latent topics…
Saranzaya Magsarjav, Melissa Humphries, Jonathan Tuke, Lewis Mitchell
Topic modelling in Natural Language Processing uncovers hidden topics in large, unlabelled text datasets. It is widely applied in fields such as information retrieval, content summarisation, and trend analysis across various disciplines. However, probabilistic topic models can produce different results when rerun due…
Filippo Valle, Matteo Osella, Michele Caselle
The integration of transcriptional data with other layers of information, such as the post-transcriptional regulation mediated by microRNAs, can be crucial to identify the driver genes and the subtypes of complex and heterogeneous diseases such as cancer. This paper presents an approach based on topic modeling to…
Filippo Valle, Michele Caselle, Matteo Osella
The availability of high-dimensional transcriptomic datasets is increasing at a tremendous pace, together with the need for suitable computational tools. Clustering and dimensionality reduction methods are popular go-to methods to identify basic structures in these datasets. At the same time, different topic modeling…
Preethi K. Periyakoil, Melanie H. Smith, Meghana Kshirsagar, Daniel Ramirez + 5 more
Single-cell RNA sequencing studies have revealed the heterogeneity of cell states present in the rheumatoid arthritis (RA) synovium. However, it remains unclear how these cell types interact with one another in situ and how synovial microenvironments shape observed cell states. Here, we use spatial transcriptomics (ST)…
Marzieh Khodaei, Scott V. Edwards, Peter Beerli
Methods for rapidly inferring the evolutionary history of species or populations with genome-wide data are progressing, but computational constraints still limit our abilities in this area. We developed an alignment-free method to infer genome-wide phylogenies and implemented it in the Python package TopicContml. The…
Authors not listed
Protein-ligand interaction prediction with proteochemometric (PCM) models can provide valuable insights during early drug discovery and chemical safety assessment. These models have benefitted from the large amount of data available in bioactivity databases. However, an issue that is often overlooked when using this…
Julian Ivanov, Alan Lipkus, Haitao Chen, Chris Aultman + 3 more
A novel bibliometric methodology based on natural language data processing for identifying emerging topics in science is presented. Along with the usual practice of data collection and preprocessing, our method includes a natural language processing (NLP) technique and an innovative mathematical function data…