24 papers · ranked by Valyu relevance
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen + 9 more
'Zhenhua Liu' 'Yehui Tang' 'An Xiao' 'Chunjing Xu' 'Yixing Xu' 'Zhaohui Yang' 'Yiman Zhang' 'Dacheng Tao'] Abstract—Transformer, first applied to the field of natural language processing, is a type of deep neural network mainly based on the self-attention mechanism. Thanks to its strong representation capabilities…
Oumaima Moutik, Hiba Sekkat, Smail Tigani, Abdellah Chehri + 4 more
'Rachid Saadane' 'Taha Ait Tchakoucht' 'Anand Paul' 'Loris Nanni'] Understanding actions in videos remains a significant challenge in computer vision, which has been the subject of several pieces of research in the last decades. Convolutional neural networks (CNN) are a significant component of this topic and play a…
Zheng Jiang, Liang Chen
The most popular test for pneumonia, a serious health threat to children, is chest X-ray imaging. However, the diagnosis of pneumonia relies on the expertise of experienced radiologists, and the scarcity of medical resources has forced us to conduct research on CAD (computer-aided diagnosis). In this study, we propose…
Bofan Song, Dharma Raj KC, Rubin Yuchan Yang, Shaobai Li + 3 more
'Chicheng Zhang' 'Rongguang Liang' 'Sam Payabvash'] Simple Summary Transformer models, originally successful in natural language processing, have found application in computer vision, demonstrating promising results in tasks related to cancer image analysis. Despite being one of the prevalent and swiftly spreading…
Gelan Ayana, Kokeb Dese, Yisak Dereje, Yonas Kebede + 7 more
'Dechassa Amdissa' 'Nahimiya Husen' 'Fikadu Mulugeta' 'Bontu Habtamu' 'Se-Woon Choe' 'Filippo Pesapane'] Breast mass identification is a crucial procedure during mammogram-based early breast cancer diagnosis. However, it is difficult to determine whether a breast lump is benign or cancerous at early stages.…
Manal Darwish, Mohamad Ziad Altabel, Rahib H. Abiyev, Malek Makki
One of the most common types of cancer among in women is cervical cancer. Incidence and fatality rates are steadily rising, particularly in developing nations, due to a lack of screening facilities, experienced specialists, and public awareness. Visual inspection is used to screen for cervical cancer after the…
Chongwen Wang, Zicheng Wang
Facial action unit (AU) detection is an important task in affective computing and has attracted extensive attention in the field of computer vision and artificial intelligence. Previous studies for AU detection usually encode complex regional feature representations with manually defined facial landmarks and learn to…
Yuda Bi, Anees Abrol, Zening Fu, Vince D. Calhoun
Deep learning models, despite their potential for increasing our understanding of intricate neuroimaging data, can be hampered by challenges related to interpretability. Multimodal neuroimaging appears to be a promising approach that allows us to extract supplementary information from various imaging modalities. It’s…
Brian Kenji Iwana, Akihiro Kusuda
Transformers are popular neural network models that use layers of self-attention and fully-connected nodes with embedded tokens. Vision Transformers (ViT) adapt transformers for image recognition tasks. In order to do this, the images are split into patches and used as tokens. One issue with ViT is the lack of…
Mingbao Lin, Mengzhao Chen, Yuxin Zhang, Yunhang Shen + 1 more
We attempt to reduce the computational costs in vision transformers (ViTs), which increase quadratically in the token number. We present a novel training paradigm that trains only one ViT model at a time, but is capable of providing improved image recognition performance with various computational costs. Here, the…
Authors not listed
Predicting protein-ligand binding affinity from three-dimensional (3D) structural data is a central task in structure-based drug discovery, yet it remains challenging due to limited data availability, structural complexity, and the sparse nature of 3D molecular representations. In this study, we investigate the…
Zilong Huang, Youcheng Ben, Guozhong Luo, Pei Cheng + 2 more
'Bin Fu'] Very recently, Window-based Transformers, which computed self-attention within non-overlapping local windows, demonstrated promising results on image classification, semantic segmentation, and object detection. However, less study has been devoted to the cross-window connection which is the key element to…
Qiang Yao, Chengyin Li, Prashant Khanduri, Dongxiao Zhu
Vision Transformers (ViTs) have become prominent models for solving various vision tasks. However, the interpretability of ViTs has not kept pace with their promising performance. While there has been a surge of interest in developing post hoc solutions to explain ViTs' outputs, these methods do not generalize to…
Authors not listed
Determining complete atomic structures directly from microscopy images remains a longstanding challenge in materials science. MicroscopyGPT is a vision-language model (VLM) that leverages multimodal generative pre-trained transformers to predict full atomic configurations including lattice parameters, element types…
Christoph Blattgerste, Tanzina Ferdous, Ayk Jessen, Maximilian Legnar + 4 more
In pathology, reconstructing adjacent tissue parts enables an overview of the macro environment of objects like tumors. Especially, malignoma are of interest to verify invasion and resection margins, as patients with positive margins face a higher mortality risk. Reassembling image fragments is widely used in other…
Matteo Farina, Pietro Zamberlan, Arno Onken, Ulisse Ferrari
For datasets with thousands of neurons and images, vision transformers have proven successful at predicting neural responses to stimuli. However, they are expected to underperform in low-data regimes, where CNNs and Gaussian processes are considered more effective. We ask whether transformers can be made competitive…
Alex Lavaee, Arkash Jain, Gustavo Scanavachi Moreira Campos, Jose Inacio Costa-Filho + 2 more
Quantitative, time-resolved 3D fluorescence microscopy can reveal complex cellular dynamics in living cells and tissues. Broader use remains limited by the difficulty of identifying, segmenting, and tracking objects of different size and shape in crowded intracellular environments in low-contrast, anisotropic…
Nikhil J. Dhinagar, Saket S. Ozarkar, Ketaki U. Buwa, Sophia I. Thomopoulos + 10 more
Recent innovations in artificial intelligence (AI) have increasingly focused on large-scale foundational models that are more general purpose in contrast to conventional models trained to perform specialized tasks. Transformer-based architectures have become the standard backbone in foundation models across data…
Chun Hung How, Jagath C. Rajapakse
An image can be seen as a long sequence of tokens with spatial structure hence an image segmentation task can be treated as a sequence-to-sequence task. Existing attention-based segmentation models have incorporated an up-sampling module at the pixel decoder, such design assumed the backbone to have multiple feature…
Kohulan Rajan, Henning Otto Brinkhaus, M. Isabel Agea, Achim Zielesny + 1 more
The number of publications describing chemical structures has increased steadily over the last decades. However, the majority of published chemical information is currently not available in machine-readable form in public databases. It remains a challenge to automate the process of information extraction in a way that…
Jonathan Skaza, Shravan Murlidaran, Apurv Varshney, Ziqi Wen + 3 more
Efforts to restore vision via neural implants have outpaced the ability to predict what users will perceive, leaving patients and clinicians without reliable tools for surgical planning or device selection. To bridge this critical gap, we introduce a computational virtual patient (CVP) pipeline that integrates…
Emma Tysinger, Brajesh Rai, Anton Sinitskiy
Meaningful exploration of the chemical space of druglike molecules in drug design is a highly challenging task due to a combinatorial explosion of possible modifications of molecules. In this work, we address this problem with transformer models, a type of machine learning (ML) model, with recent demonstrated success…
Nicole Han, Sudhanshu Srivastava, Aiwen Xu, Devi Klein + 1 more
Most retinal implants are equipped with an external VPU that is capable of applying simple image processing techniques to the video feed in real time. In the near future, these techniques may include deep learning–based algorithms aimed at improving a patient's scene understanding. Based on this premise, researchers…
Heeseung Lee, Daeho Kim, Heyin Lee, Namyoung Gwak + 6 more
- 1. Computational Science Research Center, Korea Institute of Science and Technology, Seoul 02792, Republic of Korea - 2. Department of Materials Science and Engineering, Korea University, 145 Anam-ro, Seoul 02841, Republic of Korea - 3. Department of Chemical and Biological Engineering, Korea University, Seoul 02841…