29 papers · ranked by Valyu relevance
Yile Wang, Zhanyu Shen, Hui Huang
Semantic text representation is a fundamental task in the field of natural language processing. Existing text embedding (e.g., Sim-CSE and LLM2Vec) have demonstrated excellent performance, but the values of each dimension are difficult to trace and interpret. Bag-of-words, as classic sparse interpretable embeddings…
Goran Mitrov, Boris Stanoev, Vladimir Trajkovik, Biljana Risteska Stojkoska + 4 more
Background The rapid expansion of digital data poses a unique challenge for retrieving relevant and insightful information efficiently. In particular, the increasing volume of scientific publications has made literature reviews time-consuming. The emergence of large language models (LLMs) offers new opportunities to…
John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. Rush
'Alexander M. Rush'] How much private information do text embeddings reveal about the original text? We investigate the problem of embedding inversion, reconstructing the full text represented in dense text embeddings. We frame the problem as controlled generation: generating text that, when reembedded, is close to a…
Kevin Durrheim, Maria Schuld, Martin Mafunda, Sindisiwe Mazibuko
Word embeddings provide quantitative representations of word semantics and the associations between word meanings in text data, including in large repositories in media and social media archives. This article introduces social psychologists to word embedding research via a consideration of bias analysis, a topic of…
Ryoma Sato
Word embeddings are one of the most fundamental technologies used in natural language processing. Existing word embeddings are high-dimensional and consume considerable computational resources. In this study, we propose WORDTOUR, unsupervised onedimensional word embeddings. To achieve the challenging goal, we propose a…
Dominykas Seputis, Yongkang Li, Karsten Langerak, Serghei Mihailov
Text embeddings are fundamental to many natural language processing (NLP) tasks, extensively applied in domains such as recommendation systems and information retrieval (IR). Traditionally, transmitting embeddings instead of raw text has been seen as privacy-preserving. However, recent methods such as Vec2Text…
Charles Zhang, Benji Peng, Xintian Sun, Qian Niu + 11 more
and Future Directions For Large Language Models Authors: ['Charles Zhang' 'Benji Peng' 'Xintian Sun' 'Qian Niu' 'Junyu Liu' 'Keyu Chen' 'Ming Li' 'Peiyong Feng' 'Ziqian Bi' 'Ming Liu' 'Yichao Zhang' 'Fei Cheng' 'Caitlyn Heqi Yin' 'Lingzhi Yan' 'Tianyang Wang'] Abstract—Word embeddings and language models have…
Eunji Kim, Kyuhong Shim, Simyung Chang, Sungroh Yoon
Embeddings in CLIP Authors: ['Eunji Kim' 'Kyuhong Shim' 'Simyung Chang' 'Sungroh Yoon'] A text encoder within Vision-Language Models (VLMs) like CLIP plays a crucial role in translating textual input into an embedding space shared with images, thereby facilitating the interpretative analysis of vision tasks through…
Kabane, Siyaxolisa
We investigate the generalization properties of dense text embeddings when the embedding backbone is a large language model (LLM) versus when it is a non-LLM encoder, and we study the extent to which spherical linear interpolation (SLERP) model-merging mitigates over-specialization introduced by task-specific…
Yaser A. Al-Lahham, Sattam Almatarneh, Kaznah Alshammari, Mutasem Al-Smadi
Word embedding enhances pseudo-relevance feedback query expansion (PRFQE), but training word embedding models takes a long time and is applied to large datasets. Moreover, the Arabic language, which has rich morphology, dialectal variations, and a lack of high-quality linguistic resources, training embedding models…
Wazib Ansar, Saptarsi Goswami, Amlan Chakrabarti
Evaluation using Transformers Authors: ['Wazib Ansar' 'Saptarsi Goswami' 'Amlan Chakrabarti'] One of the principal objectives of Natural Language Processing (NLP) is to generate meaningful representations from text. Improving the informativeness of the representations has led to a tremendous rise in the dimensionality…
Hassan Shahmohammadi, Maria Heitmeier, Elnaz Shafaei-Bajestan, Hendrik P. A. Lensch + 1 more
'Hendrik P. A. Lensch' 'R. Harald Baayen'] Grounding language in vision is an active field of research seeking to construct cognitively plausible word and sentence representations by incorporating perceptual knowledge from vision into text-based representations. Despite many attempts at language grounding, achieving an…
Long Qian, Xin Lu, Parvez Haris, Jianyong Zhu + 2 more
Clinical trials are crucial for drug development, but they require significant time and financial resources. Additionally, uncertainties may arise during these trials concerning their results due to concerns surrounding effectiveness, safety, or the enrollment of participants. If robust AI (artificial intelligence)…
Enock Niyonkuru, Mauricio Soto Gomez, Elena Casiraghi, Stephan Antogiovanni + 4 more
Concept embeddings are low-dimensional vector representations of concepts such as MeSH:D009203 (Myocardial Infarction), whose similarity in the embedded vector space reflects their semantic similarity. Here, we test the hypothesis that non-biomedical concept synonym replacement can improve the quality of biomedical…
Ayu Pertiwi, Azhari Azhari, Sri Mulyana, Bilal Alatas
Background Topic modeling approaches, such as latent Dirichlet allocation (LDA) and its successor, the dynamic topic model (DTM), are widely used to identify specific topics by extracting words with similar frequencies from documents. However, these topics often require manual interpretation, which poses challenges in…
Md. Aslam Parwez, Mohd. Fazil, Muhammad Arif, Md Tabrez Nafis + 1 more
'Md. Rabiul Auwul'] Due to the increasing use of information technologies by biomedical experts, researchers, public health agencies, and healthcare professionals, a large number of scientific literatures, clinical notes, and other structured and unstructured text resources are rapidly increasing and being stored in…
Sonia Maria Krißmer, Jonatan Menger, Johan Rollin, Tanja Vogel + 2 more
Pre-trained language models promise to enrich analyses of single-cell data with additional layers of information leveraging large text corpora. Yet, it is still unclear how to achieve optimal alignment with the primary quantitative single-cell data. To address this, we construct text-based training datasets from both…
Naima Oubenali, Sabrina Messaoud, Alexandre Filiot, Antoine Lamer + 1 more
'Paul Andrey'] Background Analyzing the unstructured textual data contained in electronic health records (EHRs) has always been a challenging task. Word embedding methods have become an essential foundation for neural network-based approaches in natural language processing (NLP), to learn dense and low-dimensional word…
S. S. Ho, R. E. Mills
The inundating rate of scientific publishing means every researcher will miss new discoveries from overwhelming saturation. To address this limitation, we employ natural language processing to overcome human limitations in reading, curation, and knowledge synthesis, with domain-specific applications to genetics and…
Hasan M. Sayeed, Sterling G. Baird, Taylor D. Sparks
Capturing structure-property relationships of materials for property prediction using machine learning requires the representation or featurization of the structural aspects of materials at different levels, including atomic, crystal, and microscales. While crystal structure-based modeling techniques are effective for…
Young Su Ko, Jonathan Parkinson, Wei Wang
Protein language models (pLMs) have traditionally been trained in an unsupervised manner using large protein sequence databases with an autoregressive or masked-language modeling training paradigm. Recent methods have attempted to enhance pLMs by integrating additional information, in the form of text, which are…
Harshita Sahni, Xin Chen, Trilce Estrada
Protein language models (PLMs) generate rich, layer-wise embeddings that capture diverse biological information but are expensive in terms of storage and computation at scale. In this work, we propose a compact surrogate representation for PLM embeddings across transformer layers using low-dimensional PCA projections…
Keisuke Ozawa, Teppei Suzuki, Shunsuke Tonogai, Tomoya Itakura
Developing foundation models for materials science has attracted attention. However, there is a lack of studies on inorganic materials due to the difficulty in the comprehensive representation of geometric concepts composing crystals: local atomic environments, their connections, and the global symmetries. We present a…
Saber Hafezqorani, Ka Ming Nip, Inanc Birol
Enabled by the explosion of data and substantial increase in computational power, deep learning has transformed fields such as computer vision and natural language processing (NLP) and it has become a successful method to be applied to many transcriptomic analysis tasks. A core advantage of deep learning is its…
Keisuke Ozawa, Teppei Suzuki, Shunsuke Tonogai, Tomoya Itakura
Developing foundation models for materials science has attracted attention. However, there is a lack of studies on inorganic materials due to the difficulty in the comprehensive representation of geometric concepts composing crystals: local atomic environments, their connections, and the global symmetries. We present a…
Nathan Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi + 3 more
Massive scale, both in terms of data availability and computation, enables significant breakthroughs in key application areas of deep learning such as natural language processing (NLP) and computer vision. There is emerging evidence that scale may be a key ingredient in scientific deep learning, but the importance of…
Amer El-Samman, Stijn De Baerdemacker
In deep learning methods, especially in the context of chemistry, there is an increasing urgency to uncover the hidden learning mechanisms often dubbed as ``black box." In this work, we show that graph models built on computational chemical data behave similar to natural language processing (NLP) models built on text…
Authors not listed
Compound similarity is fundamental to various cheminformatics analyses, particularly in the drug discovery industry, where the structure-activity principle is central to medicinal chemistry. Historically, binary fingerprints combined with Tanimoto and “Tanimoto-related metrics” (such as Dice, Sørensen–Dice, and…
A. Sina Booeshaghi, Aaron Streets
Large language models excel at text extraction, but they sometimes hallucinate. A simple way to avoid hallucinations is to remove any extracted text that does not appear in the original source. This is easy when the extracted text is contiguous (findable with exact string matching), but much harder when it is…