31 papers · ranked by Valyu relevance
Yile Wang, Zhanyu Shen, Hui Huang
Semantic text representation is a fundamental task in the field of natural language processing. Existing text embedding (e.g., Sim-CSE and LLM2Vec) have demonstrated excellent performance, but the values of each dimension are difficult to trace and interpret. Bag-of-words, as classic sparse interpretable embeddings…
Raúl Gómez, Lluís Gómez, Jaume Gibert, Dìmosthenis Karatzas
Self-Supervised learning from multimodal image and text data allows deep neural networks to learn powerful features with no need of human annotated data. Web and Social Media platforms provide a virtually unlimited amount of this multimodal data. In this work we propose to exploit this free available data to learn a…
Goran Mitrov, Boris Stanoev, Vladimir Trajkovik, Biljana Risteska Stojkoska + 4 more
Background The rapid expansion of digital data poses a unique challenge for retrieving relevant and insightful information efficiently. In particular, the increasing volume of scientific publications has made literature reviews time-consuming. The emergence of large language models (LLMs) offers new opportunities to…
John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. Rush
'Alexander M. Rush'] How much private information do text embeddings reveal about the original text? We investigate the problem of embedding inversion, reconstructing the full text represented in dense text embeddings. We frame the problem as controlled generation: generating text that, when reembedded, is close to a…
Kevin Durrheim, Maria Schuld, Martin Mafunda, Sindisiwe Mazibuko
Word embeddings provide quantitative representations of word semantics and the associations between word meanings in text data, including in large repositories in media and social media archives. This article introduces social psychologists to word embedding research via a consideration of bias analysis, a topic of…
Dominykas Seputis, Yongkang Li, Karsten Langerak, Serghei Mihailov
Text embeddings are fundamental to many natural language processing (NLP) tasks, extensively applied in domains such as recommendation systems and information retrieval (IR). Traditionally, transmitting embeddings instead of raw text has been seen as privacy-preserving. However, recent methods such as Vec2Text…
Yaser A. Al-Lahham, Sattam Almatarneh, Kaznah Alshammari, Mutasem Al-Smadi
Word embedding enhances pseudo-relevance feedback query expansion (PRFQE), but training word embedding models takes a long time and is applied to large datasets. Moreover, the Arabic language, which has rich morphology, dialectal variations, and a lack of high-quality linguistic resources, training embedding models…
Ilias Maglogiannis, Lazaros Iliadis, Elias Pimenidis
Word Embeddings are used widely in multiple Natural Language Processing (NLP) applications. They are coordinates associated with each word in a dictionary, inferred from statistical properties of these words in a large corpus. In this paper we introduce the notion of "concept" as a list of words that have shared…
Size Bi, Xiao Liang, Ting-lei Huang
Word embedding, a lexical vector representation generated via the neural linguistic model (NLM), is empirically demonstrated to be appropriate for improvement of the performance of traditional language model. However, the supreme dimensionality that is inherent in NLM contributes to the problems of hyperparameters and…
Ryoma Sato
Word embeddings are one of the most fundamental technologies used in natural language processing. Existing word embeddings are high-dimensional and consume considerable computational resources. In this study, we propose WORDTOUR, unsupervised onedimensional word embeddings. To achieve the challenging goal, we propose a…
Hassan Shahmohammadi, Maria Heitmeier, Elnaz Shafaei-Bajestan, Hendrik P. A. Lensch + 1 more
'Hendrik P. A. Lensch' 'R. Harald Baayen'] Grounding language in vision is an active field of research seeking to construct cognitively plausible word and sentence representations by incorporating perceptual knowledge from vision into text-based representations. Despite many attempts at language grounding, achieving an…
Mohammed Ibrahim, Susan Gauch, Tyler Gerth, Brandon C. Cox
Word vector representations open up new opportunities to extract useful information from unstructured text. Defining a word as a vector made it easy for the machine learning algorithms to understand a text and extract information from. Word vector representations have been used in many applications such word synonyms…
Long Qian, Xin Lu, Parvez Haris, Jianyong Zhu + 2 more
Clinical trials are crucial for drug development, but they require significant time and financial resources. Additionally, uncertainties may arise during these trials concerning their results due to concerns surrounding effectiveness, safety, or the enrollment of participants. If robust AI (artificial intelligence)…
Enock Niyonkuru, Mauricio Soto Gomez, Elena Casiraghi, Stephan Antogiovanni + 4 more
Concept embeddings are low-dimensional vector representations of concepts such as MeSH:D009203 (Myocardial Infarction), whose similarity in the embedded vector space reflects their semantic similarity. Here, we test the hypothesis that non-biomedical concept synonym replacement can improve the quality of biomedical…
Mario M. Kubek, Shiraj Pokharel, Thomas Böhme, Emma L. McDaniel + 2 more
'Herwig Unger' 'Armin R. Mikler'] Abstract. This article introduces a novel and fast method for refining pre-trained static word or, more generally, token embeddings. By incorporating the embeddings of neighboring tokens in text corpora, it continuously updates the representation of each token, including those without…
Md. Aslam Parwez, Mohd. Fazil, Muhammad Arif, Md Tabrez Nafis + 1 more
'Md. Rabiul Auwul'] Due to the increasing use of information technologies by biomedical experts, researchers, public health agencies, and healthcare professionals, a large number of scientific literatures, clinical notes, and other structured and unstructured text resources are rapidly increasing and being stored in…
Wei Zhuo, Qianyi Zhan, Yuan Liu, Zhenping Xie + 1 more
Network embedding (NE), which maps nodes into a low-dimensional latent Euclidean space to represent effective features of each node in the network, has obtained considerable attention in recent years. Many popular NE methods, such as DeepWalk, Node2vec, and LINE, are capable of handling homogeneous networks. However…
Sonia Maria Krißmer, Jonatan Menger, Johan Rollin, Tanja Vogel + 2 more
Pre-trained language models promise to enrich analyses of single-cell data with additional layers of information leveraging large text corpora. Yet, it is still unclear how to achieve optimal alignment with the primary quantitative single-cell data. To address this, we construct text-based training datasets from both…
S. S. Ho, R. E. Mills
The inundating rate of scientific publishing means every researcher will miss new discoveries from overwhelming saturation. To address this limitation, we employ natural language processing to overcome human limitations in reading, curation, and knowledge synthesis, with domain-specific applications to genetics and…
Hasan M. Sayeed, Sterling G. Baird, Taylor D. Sparks
Capturing structure-property relationships of materials for property prediction using machine learning requires the representation or featurization of the structural aspects of materials at different levels, including atomic, crystal, and microscales. While crystal structure-based modeling techniques are effective for…
Harshita Sahni, Xin Chen, Trilce Estrada
Protein language models (PLMs) generate rich, layer-wise embeddings that capture diverse biological information but are expensive in terms of storage and computation at scale. In this work, we propose a compact surrogate representation for PLM embeddings across transformer layers using low-dimensional PCA projections…
Young Su Ko, Jonathan Parkinson, Wei Wang
Protein language models (pLMs) have traditionally been trained in an unsupervised manner using large protein sequence databases with an autoregressive or masked-language modeling training paradigm. Recent methods have attempted to enhance pLMs by integrating additional information, in the form of text, which are…
Saber Hafezqorani, Ka Ming Nip, Inanc Birol
Enabled by the explosion of data and substantial increase in computational power, deep learning has transformed fields such as computer vision and natural language processing (NLP) and it has become a successful method to be applied to many transcriptomic analysis tasks. A core advantage of deep learning is its…
Diogo de Jesus Soares Machado, Camilla Reginatto De Pierri, Letícia Graziela Costa Santos, Leonardo Scapin + 4 more
The large amount of existing textual data justifies the development of new text mining tools. Bioinformatics tools can be brought to Text Mining, increasing the arsenal of resources. Here, we present BIOTEXT, a package of strategies for converting natural language text into biological-like information data, providing a…
Sanjar Adilov
Generative neural networks have shown promising results in de novo drug design. Recent studies suggest that one of the efficient ways to produce novel molecules matching target properties is to model SMILES sequences using deep learning in a way similar to language modeling in natural language processing. In this…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…
Amer El-Samman, Stijn De Baerdemacker
In deep learning methods, especially in the context of chemistry, there is an increasing urgency to uncover the hidden learning mechanisms often dubbed as ``black box." In this work, we show that graph models built on computational chemical data behave similar to natural language processing (NLP) models built on text…
Authors not listed
Compound similarity is fundamental to various cheminformatics analyses, particularly in the drug discovery industry, where the structure-activity principle is central to medicinal chemistry. Historically, binary fingerprints combined with Tanimoto and “Tanimoto-related metrics” (such as Dice, Sørensen–Dice, and…
Nathan Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi + 3 more
Massive scale, both in terms of data availability and computation, enables significant breakthroughs in key application areas of deep learning such as natural language processing (NLP) and computer vision. There is emerging evidence that scale may be a key ingredient in scientific deep learning, but the importance of…
A. Sina Booeshaghi, Aaron Streets
Large language models excel at text extraction, but they sometimes hallucinate. A simple way to avoid hallucinations is to remove any extracted text that does not appear in the original source. This is easy when the extracted text is contiguous (findable with exact string matching), but much harder when it is…
Rachana Niranjan Murthy, Sai Teja Potu, Akhil Thomas, Lokesh Mishra + 2 more
Retrieving structured materials information from unstructured textual data is essential for data mining and automatically developing comprehensive ontologies. Information extraction is a complex task composed of multiple subtasks and thus often relies on systems of task-specialized language models. A foundation…