23 papers · ranked by Valyu relevance
Lin Liu, Lin Tang, Wen Dong, Shaowen Yao + 1 more
Background With the rapid accumulation of biological datasets, machine learning methods designed to automate data analysis are urgently needed. In recent years, so-called topic models that originated from the field of natural language processing have been receiving much attention in bioinformatics because of their…
Josep Basha Gutierrez, Kenta Nakai
Background Topic models are statistical algorithms which try to discover the structure of a set of documents according to the abstract topics contained in them. Here we try to apply this approach to the discovery of the structure of the transcription factor binding sites (TFBS) contained in a set of biological…
Ben Curran, Kyle Higham, Elisenda Ortiz, Demival Vasques Filho + 1 more
'Floriana Gargiulo'] Quantitative methods to describe the participation to debate of Members of Parliament and the parties they belong to are lacking. Here we propose a new approach that combines topic modeling with complex networks techniques, and use it to characterize the political discourse at the New Zealand…
He Zhao, Dinh Phung, Viet Huynh, Yuan Jin + 2 more
Topic modelling has been a successful technique for text analysis for almost twenty years. When topic modelling met deep neural networks, there emerged a new and increasingly popular research area, neural topic models, with over a hundred models developed and a wide range of applications in neural language…
Sergei Koltcov, Vera Ignatenko, Zeyd Boukhers, Steffen Staab
Topic modeling is a popular technique for clustering large collections of text documents. A variety of different types of regularization is implemented in topic modeling. In this paper, we propose a novel approach for analyzing the influence of different regularization types on results of topic modeling. Based on Renyi…
Ryan Wesslen
—Topic models are a family of statistical-based algorithms to summarize, explore and index large collections of text documents. After a decade of research led by computer scientists, topic models have spread to social science as a new generation of data-driven social scientists have searched for tools to explore large…
Katherine Redfield Chang, Xinghua Lou, Theofanis Karaletsos, Christopher Crosbie + 3 more
Using a variety of techniques including Topic Modeling, Principal Component Analysis and Bi-clustering, we explore electronic patient records in the form of unstructured clinical notes and genetic mutation test results. Our ultimate goal is to gain insight into a unique body of clinical data, specifically regarding the…
Eric Austin, Shraddha Makwana, Amine Trabelsi, Christine Largeron + 1 more
Topic modeling aims to discover latent themes in collections of text documents. It has various applications across fields such as sociology, opinion analysis, and media studies. In such areas, it is essential to have easily interpretable, diverse, and coherent topics. An efficient topic modeling technique should…
Emil Rijcken, Uzay Kaymak, Floortje Scheepers, Pablo Mosteiro + 2 more
The clinical notes in electronic health records have many possibilities for predictive tasks in text classification. The interpretability of these classification models for the clinical domain is critical for decision making. Using topic models for text classification of electronic health records for a predictive task…
Kriste Krstovski, Michael J. Kurtz, David A. Smith, Alberto Accomazzi
'Alberto Accomazzi'] Scientific publications have evolved several features for mitigating vocabulary mismatch when indexing, retrieving, and computing similarity between articles. These mitigation strategies range from simply focusing on high-value article sections, such as titles and abstracts, to assigning keywords…
Márton Kardos, Jan Kostkan, Arnault‐Quentin Vermillet, Kristoffer L. Nielbo + 1 more
'Kristoffer L. Nielbo' 'Roberta Rocca'] Topic models are useful tools for discovering latent semantic structures in large textual corpora. Topic modeling historically relied on bag-of-words representations of language. This approach makes models sensitive to the presence of stop words and noise, and does not utilize…
Kyle Seelman, Mozhi Zhang, Jordan Boyd‐Graber
Topic models are valuable for understanding extensive document collections, but they don't always identify the most relevant topics. Classical probabilistic and anchor-based topic models offer interactive versions that allow users to guide the models towards more pertinent topics. However, such interactive features…
Uttam Chauhan, Shrusti Shah, Dharati Shiroya, Dipti Solanki + 9 more
'Zeel Patel' 'Jitendra Bhatia' 'Sudeep Tanwar' 'Ravi Sharma' 'Verdes Marina' 'Maria Simona Raboaca' 'Chang Choi' 'Kiho Lim' 'Gyuho Choi'] Topic modeling is a machine learning algorithm based on statistics that follows unsupervised machine learning techniques for mapping a high-dimensional corpus to a low-dimensional…
Rebecca Danning, Zheng Tracy Ke, Rong Ma, Xihong Lin
Count data are ubiquitous across many applications in which understanding hidden patterns, or latent structure, is of interest. Topic modeling is a powerful tool for detecting latent structure in count data. However, standard topic modeling methods are often constrained by their restrictive assumptions, susceptible to…
Wesam Elshamy
Topic models are probabilistic models for discovering topical themes in collections of documents. In real world applications, these models provide us with the means of organizing what would otherwise be unstructured collections. They can help us cluster a huge collection into different topics or find a subset of the…
Filippo Valle, Michele Caselle, Matteo Osella
The availability of high-dimensional transcriptomic datasets is increasing at a tremendous pace, together with the need for suitable computational tools. Clustering and dimensionality reduction methods are popular go-to methods to identify basic structures in these datasets. At the same time, different topic modeling…
Peter Carbonetto, Kaixuan Luo, Abhishek Sarkar, Anthony Hung + 3 more
Parts-based representations, such as non-negative matrix factorization and topic modeling, have been used to identify structure from single-cell sequencing data sets, in particular structure that is not as well captured by clustering or other dimensionality reduction methods. However, interpreting the individual parts…
Avinava Dubey, Ahmed Hefny, Sinead A. Williamson, Eric P. Xing
A single, stationary topic model such as latent Dirichlet allocation is inappropriate for modeling corpora that span long time periods, as the popularity of topics is likely to change over time. A number of models that incorporate time have been proposed, but in general they either exhibit limited forms of temporal…
Filippo Valle, Matteo Osella, Michele Caselle
The integration of transcriptional data with other layers of information, such as the post-transcriptional regulation mediated by microRNAs, can be crucial to identify the driver genes and the subtypes of complex and heterogeneous diseases such as cancer. This paper presents an approach based on topic modeling to…
Authors not listed
Protein-ligand interaction prediction with proteochemometric (PCM) models can provide valuable insights during early drug discovery and chemical safety assessment. These models have benefitted from the large amount of data available in bioactivity databases. However, an issue that is often overlooked when using this…
Helle W. van den Maagdenberg, Martin Šícho, David Alencar Araripe, Sohvi Luukkonen + 9 more
Building reliable and robust quantitative structure-property relationship (QSPR) models is a challenging task. First, the experimental data needs to be obtained, analyzed and curated. Second, the number of available methods is continuously growing and evaluating different algorithms and methodologies can be arduous.…
Julian Ivanov, Alan Lipkus, Haitao Chen, Chris Aultman + 3 more
A novel bibliometric methodology based on natural language data processing for identifying emerging topics in science is presented. Along with the usual practice of data collection and preprocessing, our method includes a natural language processing (NLP) technique and an innovative mathematical function data…
Xiaotong Liu, Xingchen Liu, Xiaodong Wen
The emergence of large language models (LLMs) has spurred numerous applications across various domains, including material design. In this field, an increasing number of generative models focus on directly generating materials with desired properties more accurately, moving away from the traditional approach of…