22 papers · ranked by Valyu relevance
Jonathan Roberts, Kai Han, Samuel Albanie
Frontier LLMs are increasingly utilised across academia, society and industry. A commonly used unit for comparing models, their inputs and outputs, and estimating inference pricing is the token. In general, tokens are used as a stable currency, assumed to be broadly consistent across tokenizers and contexts, enabling…
Sachin Pawar, Manoj Apte, Kshitij Jadhav, Girish Keshav Palshikar + 1 more
Tokenization is the first step in training any Large Language Model (LLM), where the text is split into a sequence of tokens as per the model's fixed vocabulary. This tokenization in LLMs is different from the traditional tokenization in NLP where the text is split into a sequence of natural words. In LLMs, a natural…
Gül Sena Altıntaş, Malikeh Ehghaghi, Brian Lester, Fengyuan Liu + 3 more
Tokenizers provide the fundamental basis through which text is represented and processed by language models (LMs). Despite the importance of tokenization, its role in LM performance and behavior is poorly understood due to the challenge of measuring the impact of tokenization in isolation. To address this need, we…
Craig W. Schmidt, Michael Krumdick, Adam Wiemerslage, Seth Ebner + 3 more
We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference procedure. ToaST greedily splits each pretoken into a full binary tree using precomputed byte n-gram counts, independent of any vocabulary. Given a vocabulary, inference…
Amit Moryossef, Clara Meister, Pavel Stepachev, Desmond Elliott
We present UTF8Tokenizer, a minimalist bytelevel tokenizer that maps text exactly to IDs corresponding to the bytes underlying the text's UTF-8 encoding (e.g., byte \x09 is token ID 9). Unlike prior byte-level approaches (Xue et al., 2021; Pagnoni et al., 2025), our implementation never introduces out-of-range IDs…
Md Toki Tahmid, Haz Sameen Shahgir, Sazan Mahbub, Yue Dong + 1 more
Transformer-based models have achieved remarkable success in biological sequence modeling, yet their application to RNA remains constrained by sequence length limitations. Existing RNA language models often truncate inputs, discarding distal nucleotide context crucial for full-length tasks. Additionally, advanced NLP…
A. Sina Booeshaghi, Aaron Streets
Large language models excel at text extraction, but they sometimes hallucinate. A simple way to avoid hallucinations is to remove any extracted text that does not appear in the original source. This is easy when the extracted text is contiguous (findable with exact string matching), but much harder when it is…
Chuxi Xiao, Yuang Ding, Haiyang Bian, Yixin Chen + 2 more
Large language models (LLMs) can process diverse forms of information once they are represented as tokens in a shared sequence space. However, single-cell transcriptomes remain a foreign modality to LLMs because they are continuous, high-dimensional molecular profiles rather than discrete linguistic units. Here, we…
Vedant Mahangade, Matthew Mollerus, Keith A. Crandall, Ali Rahnavard
Adapting language models to genomic and metagenomic sequences presents unique challenges, particularly in tokenization and task-specific generalization. Standard methods, such as fixed-length k-mers or byte pair encoding, often fail to preserve biologically meaningful patterns essential for downstream tasks. We…
Kanishk Jain, Matthew Day, Tankut Can
Tokenization introduces representational redundancy: under a fixed token vocabulary, every byte string admits many valid token encodings, or segmentations, that decode to the same surface string. However, given a prompt, most language model tokenizers break this representational symmetry by returning a canonical…
Alice Del Vecchio, Chantriolnt-Andreas Kapourani, Abdullah M. Athar, Agnieszka Dobrowolska + 4 more
DNA language models are emerging as powerful tools for representing genomic sequences, with recent progress driven by self-supervised learning. However, performance on downstream tasks is sensitive to tokenization strategies reflecting the complex encodings in DNA, where both regulatory elements and single-nucleotide…
Ruochong Zheng, Yutian Liu, Yian Zhao, Zhiwei Nie + 5 more
Three-dimensional atomic arrangements of biomolecules are key to demystifying biological functions. The rapid expansion of accessible structural data, driven by advances in AI for science, highlights the critical challenge of efficiently modeling large-scale biomolecular structures, which are high-dimensional systems…
Ella Rannon, David Burstein
Protein language models (pLMs) typically tokenize sequences at the single-amino-acid level using a 20-residue alphabet, resulting in long input sequences and high computational cost. Sub-word tokenization methods such as Byte Pair Encoding (BPE) can reduce sequence length but are limited by the sparsity of long…
Eyal Hadad, Noia Kogman, Lina Golan, Anva Avraham + 4 more
Large Language Models (LLMs) are increasingly applied to genomic tasks, yet core challenges remain concerning tokenization, evaluation, and data scarcity. This study focuses on promoter classification and systematically evaluates four tokenization methods: non-overlapping 6-mer, overlapping 6-mer, Byte Pair Encoding…
Kenny Shao
Byte Pair Encoding (BPE) is widely used for subword tokenization, but standard BPE exposes every learned merge token to the downstream model, including tokens that mainly serve as intermediate construction units and rarely appear in the final encoded corpus. This paper proposes Pruned BPE, a post-training…
Heesup Yun, Isaac Kazuo Uyehara, Ioannis Droutsas, Earl Ranario + 3 more
Three-dimensional (3D) procedural plant architecture models have emerged as an important tool for simulation-based studies of plant structure and function, extracting plant architectural parameters from field measurements, and for generating realistic plants in computer graphics. However, measuring the architectural…
Ella Rannon, David Burstein
Advances in high-throughput sequencing have generated vast amounts of genomic sequence data, much of which remains unannotated (). The ability of language models (LMs) to learn complex patterns from large unlabeled corpora makes them well-suited to address this challenge (, , , , , ). In particular, protein language…
Natthanaphop Isaradech, Wachiranun Sirikul, Stefan Schulz, Markus Kreuzthaler + 1 more
Background Extracting accurate medication information from Thai hospital records presents challenges due to the narrative style of medical notes, which often combine Thai and English terminology. Named entity recognition (NER) serves as the foundational step for advanced clinical information extraction (IE) tasks…
Shahzad Nazir, Muhammad Asif, Shahbaz Ahmad, Hanan Aljuaid + 2 more
This era has witnessed an enormous increase in textual corpus available in digital form. Therefore, an intelligent mechanism is required to extract the essential information. This task is performed using an automatic text summarization that converts the text into a shorter form while the semantics are preserved. The…
Authors not listed
Chemical language models, such as transformers trained on SMILES strings, are increasingly used in drug design and have faced rapid growth in both model capacity and training dataset size. The impact of this scaling on practical downstream performance remains unclear, however. We systematically evaluate how model size…
Authors not listed
Modeling of chemical reactions is essential for understanding kinetic mechanisms and predicting possible outcomes of reacting systems. Quantum mechanical calculations are accurate but often prohibitively expensive. Deep learning has emerged as a faster alternative, but progress is slowed by a fragmented software…
Authors not listed
Recent advances in generative artificial intelligence have enabled in silico molecular design to become a powerful approach for exploring chemical space toward specific design goals across various domains. However, in actual design workflows, determining the appropriate generation conditions, including generative…