18 papers · ranked by Valyu relevance
Manal Darwish, Mohamad Ziad Altabel, Rahib H. Abiyev, Malek Makki
One of the most common types of cancer among in women is cervical cancer. Incidence and fatality rates are steadily rising, particularly in developing nations, due to a lack of screening facilities, experienced specialists, and public awareness. Visual inspection is used to screen for cervical cancer after the…
Renán A. Rojas-Gómez, Teck-Yian Lim, Minh N. Do, Raymond A. Yeh
For computer vision, Vision Transformers (ViTs) have become one of the go-to deep net architectures. Despite being inspired by Convolutional Neural Networks (CNNs), ViTs' output remains sensitive to small spatial shifts in the input, i.e., not shift invariant. To address this shortcoming, we introduce novel…
Shelly Sheynin, Sagie Benaim, Adam Polyak, Lior Wolf
Recent work has shown the potential of transformers for computer vision applications. An image is first partitioned into patches, which are then used as input tokens for the attention mechanism. Due to the expensive quadratic cost of the attention mechanism, either a large patch size is used, resulting in…
Run Shao, Zhaoyang Zhang, Chao Tao, Yunsheng Zhang + 2 more
Sensing Image Understanding Authors: ['Run Shao' 'Zhaoyang Zhang' 'Chao Tao' 'Yunsheng Zhang' 'Chengli Peng' 'Haifeng Li'] The paradigm shift introduced by multimodal large language models, which is based on the transformer architecture and the pretext task of "next-token prediction," has revolutionized the field of…
Alice Del Vecchio, Chantriolnt-Andreas Kapourani, Abdullah M. Athar, Agnieszka Dobrowolska + 4 more
DNA language models are emerging as powerful tools for representing genomic sequences, with recent progress driven by self-supervised learning. However, performance on downstream tasks is sensitive to tokenization strategies reflecting the complex encodings in DNA, where both regulatory elements and single-nucleotide…
Marius Aasan, Odd Kolbjørnsen, Anne Solberg, Adıń Ramıŕez Rivera
Vision Transformer (ViT) architectures traditionally employ a grid-based approach to tokenization independent of the semantic content of an image. We propose a modular superpixel tokenization strategy which decouples tokenization and feature extraction; a shift from contemporary approaches where these are treated as an…
Young Kyung Kim, J. Matías Di Martino, Guillermo Sapiro
Tokens or patches within Vision Transformers (ViT) lack essential semantic information, unlike their counterparts in natural language processing (NLP). Typically, ViT tokens are associated with rectangular image patches that lack specific semantic context, making interpretation difficult and failing to effectively…
Ling Xing, Yan, Rui, Wang + 3 more
People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle typos, distorted fonts, and various scripts effectively. Modern large language models (LLMs), however, rely on subword tokenization…
Authors not listed
Identifying molecular structure based on spectroscopic readings is a key task in a va- riety of chemical and biological applications. Common spectroscopy techniques, such as Infrared (IR) Spectroscopy and Mass Spectrometry (MS), provide detailed information on the structure of molecular compounds but nonetheless…
Haris Jabbar
Tokenization is a critical part of modern NLP pipelines. However, contemporary tokenizers for Large Language Models are based on statistical analysis of text corpora, without much consideration to the linguistic features. I propose a linguistically motivated tokenization scheme, MorphPiece, which is based partly on…
Zobia Rehman, Waqas Anwar, Usama Ijaz Bajwa, Wang Xuan + 2 more
'Zhou Chaoying' 'Randen Lee Patterson'] Text tokenization is a fundamental pre-processing step for almost all the information processing applications. This task is nontrivial for the scarce resourced languages such as Urdu, as there is inconsistent use of space between words. In this paper a morpheme matching based…
Sanjar Adilov
Generative neural networks have shown promising results in de novo drug design. Recent studies suggest that one of the efficient ways to produce novel molecules matching target properties is to model SMILES sequences using deep learning in a way similar to language modeling in natural language processing. In this…
LeAnn M. Lindsey, Nicole L. Pershing, Anisa Habib, W. Zac Stephens + 2 more
Genomic language models have recently emerged as powerful tools to decode and interpret genetic sequences. Existing genomic language models have utilized various tokenization methods including character tokenization, overlapping and non-overlapping k-mer tokenization, and byte-pair encoding, a method widely used in…
Pengzhi Huang, François Charton, Jan-Niklas M. Schmelzle, Shelby S. Darnell + 3 more
The public availability of genome datasets, such as The Human Genome Project (HGP), The 1000 Genomes Project, The Cancer Genome Atlas, and the International HapMap Project, has significantly advanced scientific research and medical understanding. Here our goal is to share such genomic information for downstream…
Authors not listed
Predicting reaction yields in synthetic chemistry remains a significant challenge. This study systematically evaluates the impact of tokenization, molecular representation, pre-training data, and adversarial training on a BERT-based model for yield prediction of Buchwald-Hartwig and Suzuki-Miyaura coupling reactions…
J. A. M. Rexie, Kumudha Raimond, Mythily Murugaaboopathy, D. Brindha + 1 more
'Henock Mulugeta'] An area of medical science, that is, gaining prominence, is DNA sequencing. Genetic mutations responsible for the disease have been detected using DNA sequencing. The research is focusing on pattern identification methodologies for dealing with DNA-sequencing problems relating to various…
Neil Barrett, Jens Weber-Jahnke
Background Tokenization is an important component of language processing yet there is no widely accepted tokenization method for English texts, including biomedical texts. Other than rule based techniques, tokenization in the biomedical domain has been regarded as a classification task. Biomedical classifier-based…
Daniil Maksymenko, Oleksii Turuta
Foundational large language models (LLMs) are deployed in multilingual environments across a range of general and narrow task domains. These models generate text token by token, making them slower and more computationally expensive for low-resource languages that are underrepresented in the tokenizer vocabulary. It…