24 papers · ranked by Valyu relevance
Leonardo Costa Ribeiro, Américo Tristão Bernardes, Heliana Mello, Ramona Bongelli
'Ramona Bongelli'] Natural Language Processing (NLP) makes use of Artificial Intelligence algorithms to extract meaningful information from unstructured texts, i.e., content that lacks metadata and cannot easily be indexed or mapped onto standard database fields. It has several applications, from sentiment analysis and…
Muhammad Zulqarnain, Ahmed Khalaf Zager Alsaedi, Rozaida Ghazali, Muhammad Ghulam Ghouse + 3 more
'Muhammad Ghulam Ghouse' 'Wareesa Sharif' 'Noor Aida Husaini' 'Abdel Hamid Soliman'] Question classification is one of the essential tasks for automatic question answering implementation in natural language processing (NLP). Recently, there have been several text-mining issues such as text classification, document…
Ariel Jaffe, Yuval Kluger, Ofir Lindenbaum, Jonathan Patsenker + 2 more
'Erez Peterfreund' 'Stefan Steinerberger'] Word2vec introduced by Mikolov et al. is a word embedding method that is widely used in natural language processing. Despite its success and frequent use, a strong theoretical justification is still lacking. The main contribution of our paper is to propose a rigorous analysis…
Shihao Ji, Nadathur Satish, Sheng Li, Pradeep Dubey
—Word2Vec is a widely used algorithm for extracting low-dimensional vector representations of words. It generated considerable excitement in the machine learning and natural language processing (NLP) communities recently due to its exceptional performance in many NLP applications such as named entity recognition…
Dhananjay Kimothi, Pravesh Biyani, James M Hogan, Akshay Soni + 1 more
Similarity-based search of sequence collections is a core task in bioinformatics, one dominated for most of the genomic era by exact and heuristic alignment-based algorithms. However, even efficient heuristics such as BLAST may not scale to the data sets now emerging, motivating a range of alignment-free alternatives…
Giovanni Di Gennaro, Amedeo Buonanno, Antonio Di Girolamo, Armando Ospedale + 2 more
'Armando Ospedale' 'F. Palmieri' 'Gianfranco Fedele'] Abstract. Word representation is fundamental in NLP tasks, because it is precisely from the coding of semantic closeness between words that it is possible to think of teaching a machine to understand text. Despite the spread of word embedding concepts, still few are…
Dat Duong, Eleazar Eskin, Jingyi Jessica Li
The Gene Ontology (GO) contains GO terms that describe biological functions of genes and proteins in the cell. A GO term contains one or two sentences describing a biological aspect. GO is used in many applications. One application is the comparison of two genes or two proteins by first comparing semantic similarity of…
Qufei Chen, Marina Sokolova
In this study, we explored application of Word2Vec and Doc2Vec for sentiment analysis of clinical discharge summaries. We applied unsupervised learning since the data sets did not have sentiment annotations. Note that unsupervised learning is a more realistic scenario than supervised learning which requires an access…
Giuseppe Sgroi, Giulia Russo, Anna Maglia, Giuseppe Catanuto + 4 more
'Peter Barry' 'Andreas Karakatsanis' 'Nicola Rocco' '' 'Francesco Pappalardo'] Background Decisions in healthcare usually rely on the goodness and completeness of data that could be coupled with heuristics to improve the decision process itself. However, this is often an incomplete process. Structured interviews…
Hitoshi Iuchi, Taro Matsutani, Keisuke Yamada, Natsuki Iwano + 5 more
Remarkable advances in high-throughput sequencing have resulted in rapid data accumulation, and analyzing biological (DNA/RNA/protein) sequences to discover new insights in biology has become more critical and challenging. To tackle this issue, the application of natural language processing (NLP) to biological sequence…
Stefan Jansen
Word and phrase tables are key inputs to machine translations, but costly to produce. New unsupervised learning methods represent words and phrases in a high-dimensional vector space, and these monolingual embeddings have been shown to encode syntactic and semantic relationships between language elements. The…
Amr Al-Khatib, Samhaa R. El-Beltagy
This work presents a new and simple approach for fine-tuning pretrained word embeddings for text classification tasks. In this approach, the class in which a term appears, acts as an additional contextual variable during the fine tuning process, and contributes to the final word vector for that term. As a result, words…
Rifat Rahman
—Word embedding or vector representation of word holds syntactical and semantic characteristics of a word which can be an informative feature for any machine learningbased models of natural language processing. There are several deep learning-based models for the vectorization of words like word2vec, fasttext, gensim…
Enock Niyonkuru, Mauricio Soto Gomez, Elena Casiraghi, Stephan Antogiovanni + 4 more
Concept embeddings are low-dimensional vector representations of concepts such as MeSH:D009203 (Myocardial Infarction), whose similarity in the embedded vector space reflects their semantic similarity. Here, we test the hypothesis that non-biomedical concept synonym replacement can improve the quality of biomedical…
Halima Alachram, Hryhorii Chereda, Tim Beißbarth, Edgar Wingender + 1 more
Biomedical and life science literature is an essential way to publish experimental results. With the rapid growth of the number of new publications, the amount of scientific knowledge represented in free text is increasing remarkably. There has been much interest in developing techniques that can extract this knowledge…
Saquib Khushhal, Abdul Majid, Syed Ali Abass, Rabia Riaz + 3 more
Word embeddings are essential to natural language processing tasks because they contain a single word’s syntactic and semantic information. Word embeddings have been developed widely for numerous spoken languages across the globe like English. The research community needs to pay more attention to the Urdu language…
Ayu Pertiwi, Azhari Azhari, Sri Mulyana, Bilal Alatas
Background Topic modeling approaches, such as latent Dirichlet allocation (LDA) and its successor, the dynamic topic model (DTM), are widely used to identify specific topics by extracting words with similar frequencies from documents. However, these topics often require manual interpretation, which poses challenges in…
Chaohao Yang
Scheduling Authors: ['Chaohao Yang'] Abstract. Distributed word representation (a.k.a. word embedding) is a key focus in natural language processing (NLP). As a highly successful word embedding model, Word2Vec offers an efficient method for learning distributed word representations on large datasets. However, Word2Vec…
Daiki Yokokawa, Kazutaka Noda, Yasutaka Yanagita, Takanori Uehara + 4 more
To determine if inter-disease distances between word embedding vectors using the picot-and-cluster strategy (PCS) are a valid quantitative representation of similar disease groups in a limited domain. Abstracts were extracted from the Ichushi-Web database and subjected to morphological analysis and training using the…
Kevin Durrheim, Maria Schuld, Martin Mafunda, Sindisiwe Mazibuko
Word embeddings provide quantitative representations of word semantics and the associations between word meanings in text data, including in large repositories in media and social media archives. This article introduces social psychologists to word embedding research via a consideration of bias analysis, a topic of…
Size Bi, Xiao Liang, Ting-lei Huang
Word embedding, a lexical vector representation generated via the neural linguistic model (NLM), is empirically demonstrated to be appropriate for improvement of the performance of traditional language model. However, the supreme dimensionality that is inherent in NLM contributes to the problems of hyperparameters and…
Hasan M. Sayeed, Sterling G. Baird, Taylor D. Sparks
Capturing structure-property relationships of materials for property prediction using machine learning requires the representation or featurization of the structural aspects of materials at different levels, including atomic, crystal, and microscales. While crystal structure-based modeling techniques are effective for…
Authors not listed
Compound similarity is fundamental to various cheminformatics analyses, particularly in the drug discovery industry, where the structure-activity principle is central to medicinal chemistry. Historically, binary fingerprints combined with Tanimoto and “Tanimoto-related metrics” (such as Dice, Sørensen–Dice, and…
Tagir Akhmetshin, Arkadii Lin, Timur Madzhidov, Alexandre Varnek
Autoencoders represent a promising technique for the inverse quantitative structure-activity relationship (QSAR) task. However, undesirable bias, such as atom ordering, affects the neighbourhood behaviour of autoencoders’ latent space and, consequently, usage of the latent vectors as variables in machine-learning…