24 papers · ranked by Valyu relevance
Daniel Campos, ChengXiang Zhai
Vector-based retrieval systems have become a common staple for academic and industrial search applications because they provide a simple and scalable way of extending the search to leverage contextual representations for documents and queries. As these vectorbased systems rely on contextual language models, their usage…
Mu Yang, Andros Tjandra, Chunxi Liu, David Zhang + 3 more
'John H. L. Hansen' 'Ozlem Kalinli'] Neural network pruning compresses automatic speech recognition (ASR) models effectively. However, in multilingual ASR, languageagnostic pruning may lead to severe performance drops on some languages because language-agnostic pruning masks may not fit all languages and discard…
Haotian Xu, Yang, Jiannan, Gao + 3 more
Large Language Models (LLMs) achieve state-of-the-art performance across a wide range of applications, but their massive scale poses significant challenges for both efficiency and interpretability. Structural pruning, which reduces model size by removing redundant computational units such as neurons, has been widely…
Chao Han, Haozhe Hu, Xiaoyu Shen
Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid performance degradation beyond an essential sparsity boundary. This work asks \emph{whether combining these two mechanisms can delay such degradation by…
Mingyang Deng, Lucas Tao, Joe Benton
Recent works have proposed that activations in language models can be modelled as sparse linear combinations of vectors corresponding to features of input text. Under this assumption, these works aimed to reconstruct feature directions using sparse coding. We develop metrics to assess the success of these sparse coding…
Jingyang Yuan, Ming Zhang
Long-context modeling has become increasingly important for large language models, enabling applications such as comprehensive document analysis, extended reasoning chains and multi-turn dialogue systems. Recent models including OpenAI’s o-series, DeepSeek-R1 and Gemini 2.5 Pro have demonstrated the value of processing…
Yuxiao Li, Eric J. Michaud, David D. Baek, Joshua Engels + 4 more
'Xiaoqing Sun' 'Max Tegmark' 'Michael L. Mayo' 'Kevin R. Pilkiewicz'] Sparse autoencoders have recently produced dictionaries of high-dimensional vectors corresponding to the universe of concepts represented by large language models. We find that this concept universe has interesting structure at three levels: (1) The…
Siddhartha Brahma, Polina Zablotskaia, David Mimno
Transformers allow attention between all pairs of tokens, but there is reason to believe that most of these connections—and their quadratic time and memory—may not be necessary. But which ones? We evaluate the impact of sparsification patterns with a series of ablation experiments. First, we compare masks based on…
Kasun Vithanage, Rukshan Wijesinghe, Alex Xavier, Dumindu Tissera + 3 more
'Sanath Jayasena' 'Subha Fernando' 'Jin Liu'] In language emergence, neural agents acquire communication skills by interacting with one another and the environment. Through these interactions, agents learn to connect or ground their observations to the messages they utter, forming a shared consensus about the meaning…
Mohammad Sadat Hossain, Md. Roqunuzzaman Sojib, Md Toki Tahmid, M Saifur Rahman
RNA language models learn representations that support structure and function prediction, but which biological concepts their hidden states encode remains unclear. Sparse autoencoders (SAEs) decompose hidden states into interpretable features, yet have not been applied to RNA language models, where byte-pair…
Aaron Maiwald, Piotr Jedryszek, Florent Draye, Bernhard Schölkopf + 2 more
While genomic language models are enabling the de novo design of entire genomes, they remain challenging to interpret, limiting their trustworthiness. Here, we show that sparse autoencoders (SAEs) trained on Nucleotide Transformer activations decompose hidden representations into interpretable biological features…
Alexey V. Orlov, Yulia V. Makus, German A. Ashniev, Natalia N. Orlova + 1 more
Foundation models trained on protein and DNA sequences are increasingly deployed for variant interpretation, drug design, and gene regulation prediction, yet their internal representations remain opaque – limiting both biological insight and trust in model-guided decisions. Existing interpretation approaches establish…
Christian Brodbeck, Thomas Hannagan, James S. Magnuson, Frédéric E. Theunissen
'Frédéric E. Theunissen'] Human speech recognition transforms a continuous acoustic signal into categorical linguistic units, by aggregating information that is distributed in time. It has been suggested that this kind of information processing may be understood through the computations of a Recurrent Neural Network…
Christian Brodbeck, Thomas Hannagan, James S. Magnuson
Human speech recognition transforms a continuous acoustic signal into categorical linguistic units, by aggregating information that is distributed in time. It has been suggested that this kind of information processing may be understood through the computations of a Recurrent Neural Network (RNN) that receives input…
Chengwei Wei, Yun-Cheng Wang, Bin Wang, C.‐C. Jay Kuo
Language modeling studies the probability distributions over strings of texts. It is one of the most fundamental tasks in natural language processing (NLP). It has been widely used in text generation, speech recognition, machine translation, etc. Conventional language models (CLMs) aim to predict the probability of…
Dongqiu Zhang, Wenkui Li
Natural Language Understanding (NLU) and Natural Language Generation (NLG) are the general methods that support machine understanding of text content. They play a very important role in the text information processing system including recommendation and question and answer systems. There are many researches in the…
Yanbo Lian, Anthony N. Burkitt
Sparse coding, predictive coding and divisive normalization have each been found to be principles that underlie the function of neural circuits in many parts of the brain, supported by substantial experimental evidence. However, the connections between these related principles are still poorly understood. In this…
Shuya Nakata, Yoshiharu Mori, Shigenori Tanaka
Ultra-large virtual chemical spaces have emerged as a valuable resource for drug discovery, providing access to billions of make-on-demand compounds with high synthetic success rates. Chemical language models can potentially accelerate the exploration of these vast spaces through direct compound generation. However…
Yanbo Lian, Anthony N. Burkitt, Boris S. Gutkin
Sparse coding, predictive coding and divisive normalization have each been found to be principles that underlie the function of neural circuits in many parts of the brain, supported by substantial experimental evidence. However, the connections between these related principles are still poorly understood. Sparse coding…
Authors not listed
Predicting molecular properties is a key challenge in drug discovery. Machine learning models, especially those based on transformer architectures, are increasingly used to make these predictions from chemical structures. Inspired by recent progress in natural language processing, many studies have adopted encoder-only…
Sanjar Adilov
Machine learning models for molecular-property prediction typically work with molecular representations in the form of fingerprints, descriptors, or graphs. In case of fingerprints and descriptors, molecular representations usually comprise thousands of features, which causes the curse of dimensionality for many…
Wenzhe Yang
In this paper, we discuss how pure mathematics and theoretical physics can be applied to the study of language models. Using set theory and analysis, we formulate mathematically rigorous definitions of language models, and introduce the concept of the moduli space of distributions for a language model. We formulate a…
Yaqing Su, Lucy J. MacGregor, Itsaso Olasagasti, Anne-Lise Giraud
Understanding speech requires mapping fleeting and often ambiguous soundwaves to meaning. While humans are known to exploit their capacity to contextualize to facilitate this process, how internal knowledge is deployed on-line remains an open question. Here, we present a model that extracts multiple levels of…
Dimitris Gkoumas, Maria Liakata
The intersection of chemistry and Artificial Intelligence (AI) is an active area of research focused on accelerating scientific discovery. While using large language models (LLMs) with scientific modalities has shown potential, there are significant challenges to address, such as improving training efficiency and…