15 papers · ranked by Valyu relevance
Vasudeva Raju Sangaraju, Bharath Kumar Bolla, Deepak Kumar Nayak, Jyothsna Kh
'Jyothsna Kh'] Abstract—Customers' reviews and comments are important for businesses to understand users' sentiment about the products and services. However, this data needs to be analyzed to assess the sentiment associated with topics/aspects to provide efficient customer assistance. LDA and LSA fail to capture the…
Xiaobao Wu, T. Q. Nguyen, Anh Tuan Luu
Topic models have been prevalent for decades to discover latent topics and infer topic proportions of documents in an unsupervised fashion. They have been widely used in various applications like text analysis and context recommendation. Recently, the rise of neural networks has facilitated the emergence of a new…
Dominic B. Dayta, Erniel B. Barrios
Legacy procedures for topic modelling have generally suffered problems of overfitting and a weakness towards reconstructing sparse topic structures. This paper proposes SemiparTM, a two-step approach utilizing nonnegative matrix factorization and semiparametric regression in topic modeling. SemiparTM enables the…
Johannes Schneider
Pre-trained language models have led to a new state-of-the-art in many NLP tasks. However, for topic modeling, statistical generative models such as LDA are still prevalent, which do not easily allow incorporating contextual word vectors. They might yield topics that do not align very well with human judgment. In this…
Chengjie Ma, Junping Du, Yingxia Shao, Ang Li + 1 more
We provide a simple and general solution for the discovery of scarce topics in unbalanced short-text datasets, namely, a word co-occurrence network-based model CWIBTD, which can simultaneously address the sparsity and unbalance of short-text topics and attenuate the effect of occasional pairwise occurrences of words…
Diego Saldaña Ulloa
This work combines algorithms based on word embeddings, dimensionality reduction, and clustering. The objective is to obtain topics from a set of unclassified texts. The algorithm to obtain the word embeddings is the BERT model, a neural network architecture widely used in NLP tasks. Due to the high dimensionality, a…
Alex Gorbulev, Vasiliy Alekseev, Konstantin Vorontsov
Topic modelling is fundamentally a soft clustering problem (of known objects—documents, over unknown clusters—topics). That is, the task is incorrectly posed. In particular, the topic models are unstable and incomplete. All this leads to the fact that the process of finding a good topic model (repeated hyperparameter…
Márton Kardos, Jan Kostkan, Arnault‐Quentin Vermillet, Kristoffer L. Nielbo + 1 more
'Kristoffer L. Nielbo' 'Roberta Rocca'] Topic models are useful tools for discovering latent semantic structures in large textual corpora. Topic modeling historically relied on bag-of-words representations of language. This approach makes models sensitive to the presence of stop words and noise, and does not utilize…
Satyajeet Sahoo, Jhareswar Maiti, Virendra Kumar Tewari
—An important aspect of text mining involves information retrieval in form of discovery of semantic themes (topics) from documents using topic modelling. While generative topic models like Latent Dirichlet Allocation (LDA) elegantly model topics as probability distributions and are useful in identifying latent topics…
Saranzaya Magsarjav, Melissa Humphries, Jonathan Tuke, Lewis Mitchell
Topic modelling in Natural Language Processing uncovers hidden topics in large, unlabelled text datasets. It is widely applied in fields such as information retrieval, content summarisation, and trend analysis across various disciplines. However, probabilistic topic models can produce different results when rerun due…
Kostadin Cvejoski, Ramsés J. Sánchez, César Ojeda
Topic models and all their variants analyse text by learning meaningful representations through word co-occurrences. As pointed out by Williamson et al. (2010), such models implicitly assume that the probability of a topic to be active and its proportion within each document are positively correlated. This correlation…
Arik Reuter, Anton Thielmann, Christoph Weisser, Benjamin Säfken + 1 more
'Thomas Kneib'] Abstract—Topic modelling was mostly dominated by Bayesian graphical models during the last decade. With the rise of transformers in Natural Language Processing, however, several successful models that rely on straightforward clustering approaches in transformer-based embedding spaces have emerged and…
Huy Tran, Yating Liu, Claire Donnat
The probabilistic Latent Semantic Indexing model assumes that the expectation of the corpus matrix is low-rank and can be written as the product of a topic-word matrix and a word-document matrix. In this paper, we study the estimation of the topic-word matrix under the additional assumption that the ordered entries of…
Judicaël Poumay, Ashwin Ittoo
Topic models provide an efficient way of extracting insights from text and supporting decision-making. Recently, novel methods have been proposed to model topic hierarchy or temporality. Modeling temporality provides more precise topics by separating topics that are characterized by similar words but located over…
Jan Vávra, Bettina Grün, Paul Hofmarcher
The world is evolving and so is the vocabulary used to discuss topics in speech. Analysing political speech data from more than 30 years requires the use of flexible topic models to uncover the latent topics and their change in prevalence over time as well as the change in the vocabulary of the topics. We propose the…