Search · four archives
Search · four archives
27 papers · ranked by Valyu relevance
Wenfeng Feng, Guoying Sun
In this paper, we propose EDIT (Encoder-Decoder Image Transformer), a novel architecture designed to mitigate the attention sink phenomenon observed in Vision Transformer (ViT) models. Attention sink occurs when an excessive amount of attention is allocated to the [CLS] token, distorting the model's ability to…
Veeti Ahvonen, Damian Heiman, Antti Kuusisto, Miguel Moreno + 1 more
We give a novel logical characterization of encoder-decoder transformers, the foundational architecture for LLMs that also sees use in various settings that benefit from cross-attention. We study such transformers over text in the practical setting of floating-point numbers and soft-attention, characterizing them with…
Anna Langedijk, Hosein Mohebbi, Gabriele Sarti, Willem Zuidema + 1 more
'Jaap Jumelet'] In recent years, several interpretability methods have been proposed to interpret the inner workings of Transformer models at different levels of precision and complexity. In this work, we propose a simple but effective technique to analyze encoder-decoder Transformers. Our method, which we name…
Mamyrbayev Orken, Oralbekova Dina, Alimhan Keylan, Turdalykyzy Tolganay + 1 more
Today, the Transformer model, which allows parallelization and also has its own internal attention, has been widely used in the field of speech recognition. The great advantage of this architecture is the fast learning speed, and the lack of sequential operation, as with recurrent neural networks. In this work…
Ethan Ewer, Daewon Chae, Thomas Zeng, Jinkyu Kim + 1 more
Next-token prediction is conventionally done using decoder-only Transformers with causal attention, as this approach allows for efficient reuse of keys and values. What if we were not compute-limited, should we still use decoder-only Transformers? In this work, we introduce Encoder-only Next Token Prediction (ENTP). We…
Sumit Madan, Manuel Lentzen, Johannes Brandt, Daniel Rueckert + 2 more
'Martin Hofmann-Apitius' 'Holger Fröhlich'] Deep neural networks (DNN) have fundamentally revolutionized the artificial intelligence (AI) field. The transformer model is a type of DNN that was originally used for the natural language processing tasks and has since gained more and more attention for processing various…
Hatem Ibrahem, Ahmed Salem, Hyun-Soo Kang, Jiayi Ma
The latest research in computer vision highlighted the effectiveness of the vision transformers (ViT) in performing several computer vision tasks; they can efficiently understand and process the image globally unlike the convolution which processes the image locally. ViTs outperform the convolutional neural networks in…
Tien Thanh Thach, Antonio M. Scarfone
Accurate forecasting of stock market indices is crucial for investors, financial analysts, and policymakers. The integration of encoder and decoder architectures, coupled with an attention mechanism, has emerged as a powerful approach to enhance prediction accuracy. This paper presents a novel framework that leverages…
Kohulan Rajan, Henning Otto Brinkhaus, Achim Zielesny, Christoph Steinbeck
Accurate recognition of hand-drawn chemical structures is crucial for digitising hand-written chemical information found in traditional laboratory notebooks or for facilitating stylus-based structure entry on tablets or smartphones. However, the inherent variability in hand-drawn structures poses challenges for…
Qiumei Pu, Zuoxin Xi, Shuai Yin, Zhe Zhao + 1 more
Purpose Convolution operator-based neural networks have shown great success in medical image segmentation over the past decade. The U-shaped network with a codec structure is one of the most widely used models. Transformer, a technology used in natural language processing, can capture long-distance dependencies and has…
Satwik Bhattamishra, Arkil Patel, Navin Goyal
Transformers are being used extensively across several sequence modeling tasks. Significant research effort has been devoted to experimentally probe the inner workings of Transformers. However, our conceptual and theoretical understanding of their power and inherent limitations is still nascent. In particular, the…
Jesse Roberts
—In this article we prove that the general transformer neural model undergirding modern large language models (LLMs) is Turing complete under reasonable assumptions. This is the first work to directly address the Turing completeness of the underlying technology employed in GPT-x as past work has focused on the more…
Hossein Adeli, Sun Minni, Nikolaus Kriegeskorte
The Algonauts challenge [9] called on the community to provide novel solutions for predicting brain activity of humans viewing natural scenes. This report provides an overview and technical details of our submitted solution. We use a general transformer encoder-decoder model to map images to fMRI responses. The encoder…
Authors not listed
Deep generative models are transforming early-stage drug discovery, yet most current approaches are not well suited for realistic, small-data settings and often rely on simplified molecular representations such as linear strings, overlooking the inherent graph-based structure of molecules. To address this, we first…
Joseph G. Makin, David A. Moses, Edward F. Chang
A decade after the first successful attempt to decode speech directly from human brain signals, accuracy and speed remain far below that of natural speech or typing. Here we show how to achieve high accuracy from the electrocorticogram at natural-speech rates, even with few data (on the order of half an hour of spoken…
Harry Dong, Sean Donegan, Megna Shah, Yuejie Chi
Three dimensional electron back-scattered diffraction (EBSD) microscopy is a critical tool in many applications in materials science, yet its data quality can fluctuate greatly during the arduous collection process, particularly via serial-sectioning. Fortunately, 3D EBSD data is inherently sequential, opening up the…
Chen-Hsiu Huang, Ja-Ling Wu, Jun Chen
End-to-end learned image compression codecs have notably emerged in recent years. These codecs have demonstrated superiority over conventional methods, showcasing remarkable flexibility and adaptability across diverse data domains while supporting new distortion losses. Despite challenges such as computational…
Joel Ye, Chethan Pandarinath
Neural population activity is theorized to reflect an underlying dynamical structure. This structure can be accurately captured using state space models with explicit dynamics, such as those based on recurrent neural networks (RNNs). However, using recurrence to explicitly model dynamics necessitates sequential…
Marco Nicolini, Emanuele Saitto, Ruben Emilio Jimenez Franco, Emanuele Cavalleri + 7 more
We introduce Finenzyme, a Protein Language Model (PLM) that employs a multifaceted learning strategy based on transfer learning from a decoder-based Transformer, conditional learning using specific functional keywords, and fine-tuning to model specific Enzyme Commission (EC) categories. Using Finenzyme, we investigate…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Wenqi Guo, Yiyang Du, Mohamed Shehata
Physical molecular models are widely used in educational settings for teaching organic and other branches of chemistry, offering an intuitive way of understanding molecular structures. Conversely, virtual models, while less intuitive, provide additional functionalities such as the ability to retrieve molecular names…
Authors not listed
This research presents a novel approach to obstacle detection during navigation using a combination of Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks. The primary objective is to generate accurate image captions that describe the content of images, which is crucial for applications such…
Ian T. Ellwood
Transformers have revolutionized machine learning models of language and vision, but their connection with neuroscience remains tenuous. Built from attention layers, they require a mass comparison of queries and keys that is difficult to perform using traditional neural circuits. Here, we show that neurons can…
Samuel Renaud, Rachael Mansbach
Current antibacterial treatments cannot overcome the rapidly growing resistance of bacteria to antibiotic drugs, and novel treatment methods are required. One option is the development of new antimicrobial peptides (AMPs), to which bacterial resistance build-up is comparatively slow. Deep generative models have…
Yi Wang
Deciphering the non-coding language of DNA is one of the fundamental questions in genomic research. Previous bioinformatics methods often struggled to capture this complexity, especially in cases of limited data availability. Enhancers are short DNA segments that play a crucial role in biological processes, such as…
Authors not listed
Predicting protein-ligand binding affinity from three-dimensional (3D) structural data is a central task in structure-based drug discovery, yet it remains challenging due to limited data availability, structural complexity, and the sparse nature of 3D molecular representations. In this study, we investigate the…
Igor Sadalski
Single-cell foundational models have emerged as a powerful tool for learning generalizable cellular representations from large-scale data. Most models in this domain use transformer backbones, which require careful engineering of gene and expression encoding strategies, yet there is no consensus on which encoding…