15 papers · ranked by Valyu relevance
Rares Dolga, Lucas Maystre, Tudor Berariu, David Barber
Subword tokenization methods like Byte Pair Encoding (BPE) are widely used in large language models due to their balance of vocabulary compactness and representational power. However, they suffer from inefficiencies in representing rare words and require large embedding matrices. Character-level models address these…
Alice Del Vecchio, Chantriolnt-Andreas Kapourani, Abdullah M. Athar, Agnieszka Dobrowolska + 4 more
DNA language models are emerging as powerful tools for representing genomic sequences, with recent progress driven by self-supervised learning. However, performance on downstream tasks is sensitive to tokenization strategies reflecting the complex encodings in DNA, where both regulatory elements and single-nucleotide…
Sarah F. Elqersh, Amira Y. Haikal, Mahmoud M. Saafan, Noha A. Sakr
Accurate and efficient liver segmentation from computed tomography (CT) images remains a critical challenging task due to the organ’s irregular shape, variable intensity, and lies close to surrounding organs with similar appearance. In this study, we propose SLIC-Former, a superpixel-guided transformer framework for…
Shao, Wei, Zheng, Lingchao + 8 more
Long context inference scenarios have become increasingly important for large language models, yet they introduce significant computational latency. While prior research has optimized longsequence inference through operators, model architectures, and system frameworks, tokenization remains an overlooked bottleneck.…
Bomin Liu, Linjun He, Yan Zhu, Anil Yaman
Vision Transformers have demonstrated remarkable performance in image classification and structural modeling; however, fixed patch partitioning and static positional encoding often disrupt spatial continuity, thereby limiting their ability to represent rotated structures and irregular boundary regions. To address these…
Ling Xing, Yan, Rui, Wang + 3 more
People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle typos, distorted fonts, and various scripts effectively. Modern large language models (LLMs), however, rely on subword tokenization…
Craig W. Schmidt, Michael Krumdick, Adam Wiemerslage, Seth Ebner + 3 more
We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference procedure. ToaST greedily splits each pretoken into a full binary tree using precomputed byte n-gram counts, independent of any vocabulary. Given a vocabulary, inference…
Md Toki Tahmid, Haz Sameen Shahgir, Sazan Mahbub, Yue Dong + 1 more
Transformer-based models have achieved remarkable success in biological sequence modeling, yet their application to RNA remains constrained by sequence length limitations. Existing RNA language models often truncate inputs, discarding distal nucleotide context crucial for full-length tasks. Additionally, advanced NLP…
Vedant Mahangade, Matthew Mollerus, Keith A. Crandall, Ali Rahnavard
Adapting language models to genomic and metagenomic sequences presents unique challenges, particularly in tokenization and task-specific generalization. Standard methods, such as fixed-length k-mers or byte pair encoding, often fail to preserve biologically meaningful patterns essential for downstream tasks. We…
A. Sina Booeshaghi, Aaron Streets
Large language models excel at text extraction, but they sometimes hallucinate. A simple way to avoid hallucinations is to remove any extracted text that does not appear in the original source. This is easy when the extracted text is contiguous (findable with exact string matching), but much harder when it is…
Kanishk Jain, Matthew Day, Tankut Can
Tokenization introduces representational redundancy: under a fixed token vocabulary, every byte string admits many valid token encodings, or segmentations, that decode to the same surface string. However, given a prompt, most language model tokenizers break this representational symmetry by returning a canonical…
Geoffrey Churchill, Steven Skiena
Relative to English, low-resource languages suffer from substantial tokenization premiums in modern LMs, meaning that it generally requires several times as many tokens to encode a sentence in a low-resource language than to encode the analogous sentence in English. This tokenization premium results in increased API…
Eyal Hadad, Noia Kogman, Lina Golan, Anva Avraham + 4 more
Large Language Models (LLMs) are increasingly applied to genomic tasks, yet core challenges remain concerning tokenization, evaluation, and data scarcity. This study focuses on promoter classification and systematically evaluates four tokenization methods: non-overlapping 6-mer, overlapping 6-mer, Byte Pair Encoding…
Ella Rannon, David Burstein
Protein language models (pLMs) typically tokenize sequences at the single-amino-acid level using a 20-residue alphabet, resulting in long input sequences and high computational cost. Sub-word tokenization methods such as Byte Pair Encoding (BPE) can reduce sequence length but are limited by the sparsity of long…
Jinlei Han, Tuerhong Yushan, Juan Wang, Haitao Yang + 7 more
Most genomic foundation models are pretrained on independent linear assemblies and therefore do not explicitly represent population-level segment sharing or local graph connectivity. We developed TomatoPGFM, a graph-conditioned model pretrained on 54.65 Gb of sequence from 66 tomato (Solanum spp.) accessions. Sequence…