16 papers · ranked by Valyu relevance
Manal Darwish, Mohamad Ziad Altabel, Rahib H. Abiyev, Malek Makki
One of the most common types of cancer among in women is cervical cancer. Incidence and fatality rates are steadily rising, particularly in developing nations, due to a lack of screening facilities, experienced specialists, and public awareness. Visual inspection is used to screen for cervical cancer after the…
Renán A. Rojas-Gómez, Teck-Yian Lim, Minh N. Do, Raymond A. Yeh
For computer vision, Vision Transformers (ViTs) have become one of the go-to deep net architectures. Despite being inspired by Convolutional Neural Networks (CNNs), ViTs' output remains sensitive to small spatial shifts in the input, i.e., not shift invariant. To address this shortcoming, we introduce novel…
Run Shao, Zhaoyang Zhang, Chao Tao, Yunsheng Zhang + 2 more
Sensing Image Understanding Authors: ['Run Shao' 'Zhaoyang Zhang' 'Chao Tao' 'Yunsheng Zhang' 'Chengli Peng' 'Haifeng Li'] The paradigm shift introduced by multimodal large language models, which is based on the transformer architecture and the pretext task of "next-token prediction," has revolutionized the field of…
Md Farhadul Islam, Ishan Thakkar, J. Todd Hastings
Recent image classification models must balance local feature modeling, cross-window interaction, and parameter efficiency. Many high-performing architectures rely on fully trainable token-mixers, which improve representation learning but increase parameter count, optimization complexity and computational cost. We…
Alice Del Vecchio, Chantriolnt-Andreas Kapourani, Abdullah M. Athar, Agnieszka Dobrowolska + 4 more
DNA language models are emerging as powerful tools for representing genomic sequences, with recent progress driven by self-supervised learning. However, performance on downstream tasks is sensitive to tokenization strategies reflecting the complex encodings in DNA, where both regulatory elements and single-nucleotide…
Marius Aasan, Odd Kolbjørnsen, Anne Solberg, Adıń Ramıŕez Rivera
Vision Transformer (ViT) architectures traditionally employ a grid-based approach to tokenization independent of the semantic content of an image. We propose a modular superpixel tokenization strategy which decouples tokenization and feature extraction; a shift from contemporary approaches where these are treated as an…
Bomin Liu, Linjun He, Yan Zhu, Anil Yaman
Vision Transformers have demonstrated remarkable performance in image classification and structural modeling; however, fixed patch partitioning and static positional encoding often disrupt spatial continuity, thereby limiting their ability to represent rotated structures and irregular boundary regions. To address these…
Young Kyung Kim, J. Matías Di Martino, Guillermo Sapiro
Tokens or patches within Vision Transformers (ViT) lack essential semantic information, unlike their counterparts in natural language processing (NLP). Typically, ViT tokens are associated with rectangular image patches that lack specific semantic context, making interpretation difficult and failing to effectively…
Kalpeshkumar Ranipa, Wei-Ping Zhu, M. N. S. Swamy, Yunfeng Wu
Vision Transformers (ViTs), inspired by their success in natural language processing, have recently gained attention for heart sound classification (HSC). However, most of the existing studies on HSC rely on single-stream architectures, overlooking the advantages of multi-resolution features. While multi-stream…
Joyce Zhou, Elena L. Glassman, Daniel S. Weld
Scientists and science journalists, among others, often need to make sense of a large number of papers and how they compare with each other in scope, focus, findings, or any other important factors. However, with a large corpus of papers, it's cognitively demanding to pairwise compare and contrast them all with each…
Ling Xing, Yan, Rui, Wang + 3 more
People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle typos, distorted fonts, and various scripts effectively. Modern large language models (LLMs), however, rely on subword tokenization…
Authors not listed
Identifying molecular structure based on spectroscopic readings is a key task in a va- riety of chemical and biological applications. Common spectroscopy techniques, such as Infrared (IR) Spectroscopy and Mass Spectrometry (MS), provide detailed information on the structure of molecular compounds but nonetheless…
Pengzhi Huang, François Charton, Jan-Niklas M. Schmelzle, Shelby S. Darnell + 3 more
Language Models (LM) have been extensively utilized for learning DNA sequence patterns and generating synthetic sequences. In this paper, we present a novel approach for the generation of synthetic DNA data using pangenomes in combination with LM. We introduce three innovative pangenome-based tokenization schemes that…
Haris Jabbar
Tokenization is a critical part of modern NLP pipelines. However, contemporary tokenizers for Large Language Models are based on statistical analysis of text corpora, without much consideration to the linguistic features. I propose a linguistically motivated tokenization scheme, MorphPiece, which is based partly on…
Pengzhi Huang, François Charton, Jan-Niklas M. Schmelzle, Shelby S. Darnell + 3 more
The public availability of genome datasets, such as The Human Genome Project (HGP), The 1000 Genomes Project, The Cancer Genome Atlas, and the International HapMap Project, has significantly advanced scientific research and medical understanding. Here our goal is to share such genomic information for downstream…
Authors not listed
Predicting protein-ligand binding affinity from three-dimensional (3D) structural data is a central task in structure-based drug discovery, yet it remains challenging due to limited data availability, structural complexity, and the sparse nature of 3D molecular representations. In this study, we investigate the…