Search · four archives
Search · four archives
13 papers · ranked by Valyu relevance
Wenfeng Feng, Guoying Sun
In this paper, we propose EDIT (Encoder-Decoder Image Transformer), a novel architecture designed to mitigate the attention sink phenomenon observed in Vision Transformer (ViT) models. Attention sink occurs when an excessive amount of attention is allocated to the [CLS] token, distorting the model's ability to…
Veeti Ahvonen, Damian Heiman, Antti Kuusisto, Miguel Moreno + 1 more
We give a novel logical characterization of encoder-decoder transformers, the foundational architecture for LLMs that also sees use in various settings that benefit from cross-attention. We study such transformers over text in the practical setting of floating-point numbers and soft-attention, characterizing them with…
Anna Langedijk, Hosein Mohebbi, Gabriele Sarti, Willem Zuidema + 1 more
'Jaap Jumelet'] In recent years, several interpretability methods have been proposed to interpret the inner workings of Transformer models at different levels of precision and complexity. In this work, we propose a simple but effective technique to analyze encoder-decoder Transformers. Our method, which we name…
Mamyrbayev Orken, Oralbekova Dina, Alimhan Keylan, Turdalykyzy Tolganay + 1 more
Today, the Transformer model, which allows parallelization and also has its own internal attention, has been widely used in the field of speech recognition. The great advantage of this architecture is the fast learning speed, and the lack of sequential operation, as with recurrent neural networks. In this work…
Ethan Ewer, Daewon Chae, Thomas Zeng, Jinkyu Kim + 1 more
Next-token prediction is conventionally done using decoder-only Transformers with causal attention, as this approach allows for efficient reuse of keys and values. What if we were not compute-limited, should we still use decoder-only Transformers? In this work, we introduce Encoder-only Next Token Prediction (ENTP). We…
Sumit Madan, Manuel Lentzen, Johannes Brandt, Daniel Rueckert + 2 more
'Martin Hofmann-Apitius' 'Holger Fröhlich'] Deep neural networks (DNN) have fundamentally revolutionized the artificial intelligence (AI) field. The transformer model is a type of DNN that was originally used for the natural language processing tasks and has since gained more and more attention for processing various…
Hatem Ibrahem, Ahmed Salem, Hyun-Soo Kang, Jiayi Ma
The latest research in computer vision highlighted the effectiveness of the vision transformers (ViT) in performing several computer vision tasks; they can efficiently understand and process the image globally unlike the convolution which processes the image locally. ViTs outperform the convolutional neural networks in…
Tien Thanh Thach, Antonio M. Scarfone
Accurate forecasting of stock market indices is crucial for investors, financial analysts, and policymakers. The integration of encoder and decoder architectures, coupled with an attention mechanism, has emerged as a powerful approach to enhance prediction accuracy. This paper presents a novel framework that leverages…
Qiumei Pu, Zuoxin Xi, Shuai Yin, Zhe Zhao + 1 more
Purpose Convolution operator-based neural networks have shown great success in medical image segmentation over the past decade. The U-shaped network with a codec structure is one of the most widely used models. Transformer, a technology used in natural language processing, can capture long-distance dependencies and has…
Satwik Bhattamishra, Arkil Patel, Navin Goyal
Transformers are being used extensively across several sequence modeling tasks. Significant research effort has been devoted to experimentally probe the inner workings of Transformers. However, our conceptual and theoretical understanding of their power and inherent limitations is still nascent. In particular, the…
Jesse Roberts
—In this article we prove that the general transformer neural model undergirding modern large language models (LLMs) is Turing complete under reasonable assumptions. This is the first work to directly address the Turing completeness of the underlying technology employed in GPT-x as past work has focused on the more…
Harry Dong, Sean Donegan, Megna Shah, Yuejie Chi
Three dimensional electron back-scattered diffraction (EBSD) microscopy is a critical tool in many applications in materials science, yet its data quality can fluctuate greatly during the arduous collection process, particularly via serial-sectioning. Fortunately, 3D EBSD data is inherently sequential, opening up the…
Chen-Hsiu Huang, Ja-Ling Wu, Jun Chen
End-to-end learned image compression codecs have notably emerged in recent years. These codecs have demonstrated superiority over conventional methods, showcasing remarkable flexibility and adaptability across diverse data domains while supporting new distortion losses. Despite challenges such as computational…