23 papers · ranked by Valyu relevance
Craig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine + 3 more
'Omri Uzan' 'Yuval Pinter' 'Chris C. Tanner'] Tokenization is a foundational step in Natural Language Processing (NLP) tasks, bridging raw text and language models. Existing tokenization approaches like Byte-Pair Encoding (BPE) originate from the field of data compression, and it has been suggested that the…
Martin Berglund, Brink van der Merwe
In this paper, we formalize practical byte pair encoding tokenization as it is used in large language models and other NLP systems, in particular we formally define and investigate the semantics of the SentencePiece and HuggingFace tokenizers, in particular how they relate to each other, depending on how the…
LeAnn M Lindsey, Nicole L Pershing, Anisa Habib, Keith Dufault-Thompson + 5 more
Tokenization is a fundamental step in the language model preprocessing pipeline and is used to parse an input sequence into segments called tokens that represent either words, subwords, or characters. These tokens are assigned numeric values and used as inputs to a neural network in order to learn context-specific…
Neil Barrett, Jens Weber-Jahnke
Background Tokenization is an important component of language processing yet there is no widely accepted tokenization method for English texts, including biomedical texts. Other than rule based techniques, tokenization in the biomedical domain has been regarded as a classification task. Biomedical classifier-based…
Zaid Alyafeai, Maged S. Al-shaibani, Mustafa Ghaleb, Irfan Ahmad
The first step in any NLP pipeline is to split the text into individual tokens. The most obvious and straightforward approach is to use words as tokens. However, given a large text corpus, representing all the words is not efficient in terms of vocabulary size. In the literature, many tokenization algorithms have…
LeAnn M. Lindsey, Nicole L. Pershing, Anisa Habib, W. Zac Stephens + 2 more
Genomic language models have recently emerged as powerful tools to decode and interpret genetic sequences. Existing genomic language models have utilized various tokenization methods including character tokenization, overlapping and non-overlapping k-mer tokenization, and byte-pair encoding, a method widely used in…
Bell Raj Eapen
Transformer models are revolutionizing sequence analysis across various domains, from natural language processing to genomics. These models rely on tokenizers to split input sequences into manageable chunks — a straightforward task in natural language but more challenging for long DNA sequences that lack distinct…
Shahzad Nazir, Muhammad Asif, Mariam Rehman, Shahbaz Ahmad + 1 more
'Xiangjie Kong'] In text applications, pre-processing is deemed as a significant parameter to enhance the outcomes of natural language processing (NLP) chores. Text normalization and tokenization are two pivotal procedures of text pre-processing that cannot be overstated. Text normalization refers to transforming raw…
Bharath Raj S, Gaurav Suri, Vikrant Dewangan, Raghav Sonavane
Models Authors: ['Bharath Raj S' 'Gaurav Suri' 'Vikrant Dewangan' 'Raghav Sonavane'] Traditional greedy tokenization methods have been a critical step in Natural Language Processing (NLP), influencing how text is converted into tokens and directly impacting model performance. While subword tokenizers like Byte-Pair…
Zobia Rehman, Waqas Anwar, Usama Ijaz Bajwa, Wang Xuan + 2 more
'Zhou Chaoying' 'Randen Lee Patterson'] Text tokenization is a fundamental pre-processing step for almost all the information processing applications. This task is nontrivial for the scarce resourced languages such as Urdu, as there is inconsistent use of space between words. In this paper a morpheme matching based…
Pengzhi Huang, François Charton, Jan-Niklas M. Schmelzle, Shelby S. Darnell + 3 more
The public availability of genome datasets, such as The Human Genome Project (HGP), The 1000 Genomes Project, The Cancer Genome Atlas, and the International HapMap Project, has significantly advanced scientific research and medical understanding. Here our goal is to share such genomic information for downstream…
Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao + 2 more
with Model Performance Authors: ['Omer Goldman' 'Avi Caciularu' 'Matan Eyal' 'Kris Cao' 'Idan Szpektor' 'Reut Tsarfaty'] Despite it being the cornerstone of BPE, the most common tokenization algorithm, the importance of compression in the tokenization process is still unclear. In this paper, we argue for the…
Md Toki Tahmid, Haz Sameen Shahgir, Sazan Mahbub, Yue Dong + 1 more
Recent advancements in Transformer-based models have spurred interest in their use for biological sequence analysis. However, adapting models like BERT is challenging due to sequence length, often requiring truncation for proteomics and genomics tasks. Additionally, advanced tokenization and relative positional…
Md Toki Tahmid, Haz Sameen Shahgir, Sazan Mahbub, Yue Dong + 1 more
Transformer-based models have achieved remarkable success in biological sequence modeling, yet their application to RNA remains constrained by sequence length limitations. Existing RNA language models often truncate inputs, discarding distal nucleotide context crucial for full-length tasks. Additionally, advanced NLP…
Christopher Meaney, Thérèse A. Stukel, Peter C. Austin, Michael Escobar
'Michael Escobar'] Background & Objective: Biomedical text data are increasingly available for research. Tokenization is an initial step in many biomedical text mining pipelines. Tokenization is the process of parsing an input biomedical sentence (represented as a digital character sequence) into a discrete set of…
Authors not listed
Predicting reaction yields in synthetic chemistry remains a significant challenge. This study systematically evaluates the impact of tokenization, molecular representation, pre-training data, and adversarial training on a BERT-based model for yield prediction of Buchwald-Hartwig and Suzuki-Miyaura coupling reactions…
Philip A. Whittington, Gregor Bachmann, Tiago Pimentel
Tokenisation is at the heart of natural language processing (NLP) being the first step required to use a language model (LM). Given a string of characters c, a tokeniser converts it into a string of subwords s. Language models are then trained to estimate distributions over subword strings—never seeing the original…
Ella Rannon, David Burstein
Protein language models (pLMs) typically tokenize sequences at the single-amino-acid level using a 20-residue alphabet, resulting in long input sequences and high computational cost. Sub-word tokenization methods such as Byte Pair Encoding (BPE) can reduce sequence length but are limited by the sparsity of long…
Esben Bjerrum, Tobias Rastemo, Ross Irwin, Christos Kannas + 1 more
Recent years have seen a large interest in using the Simplified Molecular Input Line Entry System (SMILES) chemical language as input for deep learning architectures solving chemical tasks. Many successful applications have been demonstrated within de novo molecular design, quantitative structure-activity relationship…
Kohulan Rajan, Henning Otto Brinkhaus, M. Isabel Agea, Achim Zielesny + 1 more
The number of publications describing chemical structures has increased steadily over the last decades. However, the majority of published chemical information is currently not available in machine-readable form in public databases. It remains a challenge to automate the process of information extraction in a way that…
Authors not listed
Deep generative models are transforming early-stage drug discovery, yet most current approaches are not well suited for realistic, small-data settings and often rely on simplified molecular representations such as linear strings, overlooking the inherent graph-based structure of molecules. To address this, we first…
Sanjar Adilov
Generative neural networks have shown promising results in de novo drug design. Recent studies suggest that one of the efficient ways to produce novel molecules matching target properties is to model SMILES sequences using deep learning in a way similar to language modeling in natural language processing. In this…
Andrew Blanchard, Pei Zhang, Debsindhu Bhowmik, Kshitij Mehta + 4 more
Efficient methods for searching the chemical space of molecular compounds are needed to automate and accelerate the design of new functional molecules such as pharmaceuticals. Given the high cost in both resources and time for experimental efforts, computational approaches play a key role in guiding the selection of…