26 papers · ranked by Valyu relevance
Anil Rahate, Rahee Walambe, Sheela Ramanna, Ketan Kotecha
Multimodal deep learning systems that employ multiple modalities like text, image, audio, video, etc., are showing better performance than individual modalities (i.e., unimodal) systems. Multimodal machine learning involves multiple aspects: representation, translation, alignment, fusion, and co-learning. In the…
Kuan Liu, Yanen Li, Ning Xu, Prem Natarajan
Combining complementary information from multiple modalities is intuitively appealing for improving the performance of learning-based approaches. However, it is challenging to fully leverage different modalities due to practical challenges such as varying levels of noise and conflicts between modalities. Existing…
Jabeen Summaira, Xi Li, Amin Muhammad Shoib, Songyuan Li + 1 more
'Jabbar Abdul'] Deep Learning has implemented a wide range of applications and has become increasingly popular in recent years. The goal of multimodal deep learning is to create models that can process and link information using various modalities. Despite the extensive development made for unimodal learning, it still…
Vinitra Swamy, Malika Satayeva, Jibril Frej, Thierry Bossy + 4 more
'Thijs Vogels' 'Martin Jaggi' 'Tanja Käser' 'Mary-Anne Hartley'] | Vinitra Swamy | Malika Satayeva | Jibril Frej | Thierry Bossy | | --- | --- | --- | --- | | EPFL | EPFL | EPFL | EPFL | | vinitra.swamy@epfl.ch | malika.satayeva@epfl.ch | jibril.frej@epfl.ch | thierry.bossy@epfl.ch | | Thijs Vogels | Martin Jaggi |…
Yuda Bi, Anees Abrol, Zening Fu, Vince D. Calhoun
Deep learning models, despite their potential for increasing our understanding of intricate neuroimaging data, can be hampered by challenges related to interpretability. Multimodal neuroimaging appears to be a promising approach that allows us to extract supplementary information from various imaging modalities. It’s…
Xiang He, Dongcheng Zhao, Yang Li, Qingqun Kong + 2 more
Multimodal learning enhances the perceptual ability of intelligent systems by integrating information across sensory modalities. However, most artificial intelligence approaches still rely on static fusion schemes and do not account for the dynamic nature of multisensory integration observed in the brain. In biological…
Garam Lee, Byungkon Kang, Kwangsik Nho, Kyung-Ah Sohn + 1 more
As large amounts of heterogeneous biomedical data become available, numerous methods for integrating such datasets have been developed to extract complementary knowledge from multiple domains of sources. Recently, a deep learning approach has shown promising results in a variety of research areas. However, applying the…
Vijay John, Yasutomo Kawanishi, Stefanos Kollias
In classification tasks, such as face recognition and emotion recognition, multimodal information is used for accurate classification. Once a multimodal classification model is trained with a set of modalities, it estimates the class label by using the entire modality set. A trained classifier is typically not…
Olaide N. Oyelade, Eric Aghiomesi Irunokhai, Hui Wang
There is a wide application of deep learning technique to unimodal medical image analysis with significant classification accuracy performance observed. However, real-world diagnosis of some chronic diseases such as breast cancer often require multimodal data streams with different modalities of visual and textual…
Hakim Benkirane, Maria Vakalopoulou, David Planchard, Julien Adam + 3 more
Characterizing cancer poses a delicate challenge as it involves deciphering complex biological interactions within the tumor’s microenvironment. Histology images and molecular profiling of tumors are often available in clinical trials and can be leveraged to understand these interactions. However, despite recent…
Douwe Kiela, Édouard Grave, Armand Joulin, Tomáš Mikolov
While the incipient internet was largely text-based, the modern digital world is becoming increasingly multi-modal. Here, we examine multi-modal classification where one modality is discrete, e.g. text, and the other is continuous, e.g. visual representations transferred from a convolutional neural network. In…
Till Richter, Eric Zimmermann, James Hall, Fabian J. Theis + 4 more
The vision of a “virtual cell”—a computational model that simulates biological function across modalities and scales—has become a defining goal in computational biology. While powerful unimodal foundation models exist, the lack of large-scale paired data prohibits the joint training of multimodal approaches. This…
William C. Sleeman, Rishabh Kapoor, Preetam Ghosh
Multimodal classification research has been gaining popularity in many domains that collect more data from multiple sources including satellite imagery, biometrics, and medicine. However, the lack of consistent terminology and architectural descriptions makes it difficult to compare different existing solutions. We…
Zarif L. Azher, Anish Suvarna, Ji-Qing Chen, Ze Zhang + 4 more
Deep learning models have demonstrated the remarkable ability to infer cancer patient prognosis from molecular and anatomic pathology information. Studies in recent years have demonstrated that leveraging information from complementary multimodal data can improve prognostication, further illustrating the potential…
Charles A. Ellis, Darwin A. Carbajal, Rongen Zhang, Robyn L. Miller + 2 more
In recent years, more biomedical studies have begun to use multimodal data to improve model performance. As such, there is a need for improved multimodal explainability methods. Many studies involving multimodal explainability have used ablation approaches. Ablation requires the modification of input data, which may…
Austin Reiter, Menglin Jia, Pu-Tai Yang, Ser-Nam Lim
Many vision-related tasks benefit from reasoning over multiple modalities to leverage complementary views of data in an attempt to learn robust embedding spaces. Most deep learning-based methods rely on a late fusion technique whereby multiple feature types are encoded and concatenated and then a multi layer perceptron…
Noah Cohen Kalafut, Xiang Huang, Daifeng Wang
Single-cell multimodal datasets have measured various characteristics of individual cells, enabling a deep understanding of cellular and molecular mechanisms. However, multimodal data generation remains costly and challenging, and missing modalities happen frequently. Recently, machine learning approaches have been…
Xin Chang, Władysław Skarbek, Stefano Berretti
Emotion recognition is an important research field for human-computer interaction. Audio-video emotion recognition is now attacked with deep neural network modeling tools. In published papers, as a rule, the authors show only cases of the superiority in multi-modality over audio-only or video-only modality. However…
Junwei Liu, Xiaoping Cen, Chenxin Yi, Feng-ao Wang + 9 more
The rapid development of biological and medical examination methods has vastly expanded personal biomedical information, including molecular, cellular, image, and electronic health record datasets. Integrating this wealth of information enables precise disease diagnosis, biomarker identification, and treatment design…
Elisa Warner, Joonsang Lee, William Hsu, Tanveer Syeda-Mahmood + 3 more
'Charles E. Kahn Jr.' 'Olivier Gevaert' 'Arvind Rao'] Machine learning (ML) applications in medical artificial intelligence (AI) systems have shifted from traditional and statistical methods to increasing application of deep learning models. This survey navigates the current landscape of multimodal ML, focusing on its…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Authors not listed
Early prediction of drug-induced organ toxicity remains a major bottleneck in drug discovery and clinical pharmacotherapy. Most data-driven toxicity models behave as endpoint predictors: they output a label but provide limited transparency about why a compound is risky or which evidence channel dominated the decision.…
Authors not listed
High-entropy layered double hydroxides (HE-LDHs) have shown great potential in oxygen evolution reaction (OER) catalysis due to their tunable compositions and electronic structures. However, the synergistic effects between multiple vacancies, such as metal and oxygen vacancies, remain poorly understood and challenging…
Authors not listed
Identifying molecular structure based on spectroscopic readings is a key task in a va- riety of chemical and biological applications. Common spectroscopy techniques, such as Infrared (IR) Spectroscopy and Mass Spectrometry (MS), provide detailed information on the structure of molecular compounds but nonetheless…
Derek van Tilborg, Helena Brinkmann, Emanuele Criscuolo, Luke Rossen + 2 more
Deep learning is becoming increasingly relevant in drug discovery, from de novo design to protein structure prediction and synthesis planning. However, it is often challenged by the small data regimes typical of certain drug discovery tasks. In such scenarios, deep learning approaches – which are notoriously…
Zachary Humphreys, Xenophon Evangelopoulos, Stavros Gerolymatos, Edward O. Pyzer-Knapp + 1 more
Graph neural networks have recently met huge success in various inference tasks including materials property prediction amongst many others. Nevertheless, having an inherently locally-based representation capacity as they do, global representation of materials' structures can only only be achieved by expanding the…