26 papers · ranked by Valyu relevance
Daniel Campos, ChengXiang Zhai
Vector-based retrieval systems have become a common staple for academic and industrial search applications because they provide a simple and scalable way of extending the search to leverage contextual representations for documents and queries. As these vectorbased systems rely on contextual language models, their usage…
Mingyang Deng, Lucas Tao, Joe Benton
Recent works have proposed that activations in language models can be modelled as sparse linear combinations of vectors corresponding to features of input text. Under this assumption, these works aimed to reconstruct feature directions using sparse coding. We develop metrics to assess the success of these sparse coding…
Naomi Saphra, Adam Lopez
Concerns about interpretability, computational resources, and principled inductive priors have motivated efforts to engineer sparse neural models for NLP tasks. If sparsity is important for NLP, might well-trained neural models naturally become roughly sparse? Using the Taxi-Euclidean norm to measure sparsity, we find…
Yuxiao Li, Eric J. Michaud, David D. Baek, Joshua Engels + 4 more
'Xiaoqing Sun' 'Max Tegmark' 'Michael L. Mayo' 'Kevin R. Pilkiewicz'] Sparse autoencoders have recently produced dictionaries of high-dimensional vectors corresponding to the universe of concepts represented by large language models. We find that this concept universe has interesting structure at three levels: (1) The…
Yihong Gu, Jun Yan, Hao Zhu, Zhiyuan Liu + 4 more
'Fen Lin' 'Leyu Lin'] Most language modeling methods rely on large-scale data to statistically learn the sequential patterns of words. In this paper, we argue that words are atomic language units but not necessarily atomic semantic units. Inspired by HowNet, we use sememes, the minimum semantic units in human…
Thomas Demeester, Johannes Deleu, Fréderic Godin, Chris Develder
Inducing sparseness while training neural networks has been shown to yield models with a lower memory footprint but similar effectiveness to dense models. However, sparseness is typically induced starting from a dense model, and thus this advantage does not hold during training. We propose techniques to enforce…
Mohammad Sadat Hossain, Md. Roqunuzzaman Sojib, Md Toki Tahmid, M Saifur Rahman
RNA language models learn representations that support structure and function prediction, but which biological concepts their hidden states encode remains unclear. Sparse autoencoders (SAEs) decompose hidden states into interpretable features, yet have not been applied to RNA language models, where byte-pair…
Aaron Maiwald, Piotr Jedryszek, Florent Draye, Bernhard Schölkopf + 2 more
While genomic language models are enabling the de novo design of entire genomes, they remain challenging to interpret, limiting their trustworthiness. Here, we show that sparse autoencoders (SAEs) trained on Nucleotide Transformer activations decompose hidden representations into interpretable biological features…
Alexey V. Orlov, Yulia V. Makus, German A. Ashniev, Natalia N. Orlova + 1 more
Foundation models trained on protein and DNA sequences are increasingly deployed for variant interpretation, drug design, and gene regulation prediction, yet their internal representations remain opaque – limiting both biological insight and trust in model-guided decisions. Existing interpretation approaches establish…
Wenpeng Hu, Mengyu Wang, Bing Liu, Feng Ji + 4 more
'Dongyan Zhao' 'Jinwen Ma' 'Rui Yan'] Sparsity is regarded as a desirable property of representations, especially in terms of explanation. However, its usage has been limited due to the gap with dense representations. Most NLP research progresses in recent years are based on dense representations. Thus the desirable…
Sylvester Olubolu Orimaye, Jojo Sze-Meng Wong, Chee Piau Wong, Peipeng Liang
'Peipeng Liang'] It has been quite a challenge to diagnose Mild Cognitive Impairment due to Alzheimer’s disease (MCI) and Alzheimer-type dementia (AD-type dementia) using the currently available clinical diagnostic criteria and neuropsychological examinations. As such we propose an automated diagnostic technique using…
Yair Lakretz, Stanislas Dehaene, Jean-Rémi King
Sentence comprehension requires inferring, from a sequence of words, the structure of syntactic relationships that bind these words into a semantic representation. Our limited ability to build some specific syntactic structures, such as nested center-embedded clauses (e.g., “The dog that the cat that the mouse bit…
Christian Brodbeck, Thomas Hannagan, James S. Magnuson, Frédéric E. Theunissen
'Frédéric E. Theunissen'] Human speech recognition transforms a continuous acoustic signal into categorical linguistic units, by aggregating information that is distributed in time. It has been suggested that this kind of information processing may be understood through the computations of a Recurrent Neural Network…
Christian Brodbeck, Thomas Hannagan, James S. Magnuson
Human speech recognition transforms a continuous acoustic signal into categorical linguistic units, by aggregating information that is distributed in time. It has been suggested that this kind of information processing may be understood through the computations of a Recurrent Neural Network (RNN) that receives input…
Lea-Maria Schmitt, Julia Erb, Sarah Tune, Anna Rysop + 2 more
How can anticipatory neural processes structure the temporal unfolding of context in our natural environment? We here provide evidence for a neural coding scheme that sparsely updates contextual representations at the boundary of events and gives rise to a hierarchical, multi-layered organization of predictive language…
Yanbo Lian, Anthony N. Burkitt, Boris S. Gutkin
Sparse coding, predictive coding and divisive normalization have each been found to be principles that underlie the function of neural circuits in many parts of the brain, supported by substantial experimental evidence. However, the connections between these related principles are still poorly understood. Sparse coding…
Gabriel Barello, Adam S. Charles, Jonathan W. Pillow
The sparse coding model posits that the visual system has evolved to efficiently code natural stimuli using a sparse set of features from an overcomplete dictionary. The classic sparse coding model suffers from two key limitations, however: (1) computing the neural response to an image patch requires minimizing a…
Authors not listed
Predicting molecular properties is a key challenge in drug discovery. Machine learning models, especially those based on transformer architectures, are increasingly used to make these predictions from chemical structures. Inspired by recent progress in natural language processing, many studies have adopted encoder-only…
Shuya Nakata, Yoshiharu Mori, Shigenori Tanaka
Ultra-large virtual chemical spaces have emerged as a valuable resource for drug discovery, providing access to billions of make-on-demand compounds with high synthetic success rates. Chemical language models can potentially accelerate the exploration of these vast spaces through direct compound generation. However…
Sanjar Adilov
Generative neural networks have shown promising results in de novo drug design. Recent studies suggest that one of the efficient ways to produce novel molecules matching target properties is to model SMILES sequences using deep learning in a way similar to language modeling in natural language processing. In this…
Sanjar Adilov
Machine learning models for molecular-property prediction typically work with molecular representations in the form of fingerprints, descriptors, or graphs. In case of fingerprints and descriptors, molecular representations usually comprise thousands of features, which causes the curse of dimensionality for many…
Dimitris Gkoumas, Maria Liakata
The intersection of chemistry and Artificial Intelligence (AI) is an active area of research focused on accelerating scientific discovery. While using large language models (LLMs) with scientific modalities has shown potential, there are significant challenges to address, such as improving training efficiency and…
Yushi Sugimoto, Ryo Yoshida, Hyeonjeong Jeong, Masatoshi Koizumi + 2 more
'Jonathan R. Brennan' 'Yohei Oseki'] Title: Abstract In computational neurolinguistics, it has been demonstrated that hierarchical models such as recurrent neural network grammars (RNNGs), which jointly generate word sequences and their syntactic structures via the syntactic composition, better explained human brain…
Rachana Niranjan Murthy, Sai Teja Potu, Akhil Thomas, Lokesh Mishra + 2 more
Retrieving structured materials information from unstructured textual data is essential for data mining and automatically developing comprehensive ontologies. Information extraction is a complex task composed of multiple subtasks and thus often relies on systems of task-specialized language models. A foundation…
Ross Irwin, Spyridon Dimitriadis, Jiazhen He, Esben Bjerrum
Transformer models coupled with Simplified Molecular Line Entry System (SMILES) have recently proven to be a powerful combination for solving challenges in cheminformatics. These models, however, are often developed specifically for a single application and can be very resource-intensive to train. In this work we…
Authors not listed
Natural language processing with the help of large language models such as ChatGPT has become ubiquitous in many software applications and allows users to interact even with complex hardware or software in an intuitive way. The recent concepts of Self-Driving Labs and Material Acceleration Platforms stand to benefit…