14 papers · ranked by Valyu relevance
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski, Maciej Pióro + 8 more
'Michał Krutul' 'Szymon Antoniak' 'Kamil Ciebiera' 'Krystian Król' 'Tomasz Odrzygóźdź' 'Piotr Sankowski' 'Marek Cygan' 'Sebastian Jaszczur'] | Jakub Krajewski ∗ | Jan Ludziejewski ∗ | Kamil Adamczewski | Maciej Pioro ´ | | --- | --- | --- | --- | | University of Warsaw | University of Warsaw | IDEAS NCBR | IPPT PAN | |…
Youngseog Chung, Dhruv Malik, Jeff Schneider, Yuanzhi Li + 1 more
'Aarti Singh'] The traditional viewpoint on Sparse Mixture of Experts (MoE) models is that instead of training a single large expert, which is computationally expensive, we can train many small experts. The hope is that if the total parameter count of the small experts equals that of the singular large expert, then we…
Ning Sun, Shuxian Zou, Tianhua Tao, Sazan Mahbub + 6 more
Proteins play a fundamental role in life. Understanding the language of proteins offers significant potential for gaining mechanistic insights into biological systems and introduces new avenues for treating diseases, enhancing agriculture, and safeguarding the environment. While large protein language models (PLMs)…
James Oldfield, Markos Georgopoulos, Grigorios G. Chrysos, Christos Tzelepis + 4 more
Factorization Authors: ['James Oldfield' 'Markos Georgopoulos' 'Grigorios G. Chrysos' 'Christos Tzelepis' 'Yannis Panagakis' 'Mihalis A. Nicolaou' 'Jiankang Deng' 'Ioannis Patras'] The Mixture of Experts (MoE) paradigm provides a powerful way to decompose inscrutable dense layers into smaller, modular computations…
Shadi Zabad, Yue Li, Simon Gravel
With the increasing availability of high quality genomic data from diverse cohorts, polygenic scores (PRS) have become a mainstay of genetic analyses of complex traits and diseases. Despite their proliferation in numerous research domains, a major obstacle to wider adoption in clinical settings has been the…
Xiang Zhang, Shenbao Yu, Jie Xia, Fan Yang
Recent advancements in large-scale self-supervised pretraining have significantly improved molecular representation learning, yet challenges persist, particularly when addressing distributional shifts (e.g., under scaffold-split). Drawing inspiration from the success of Mixture-of-Experts (MoE) networks in NLP, we…
Ziwei Zhan, Wei Zhao, Yuanqing Li, Weijie Liu + 5 more
Aggregation Authors: ['Ziwei Zhan' 'Wei Zhao' 'Yuanqing Li' 'Weijie Liu' 'X.C. Zhang' 'Chee Wei Tan' 'Chuanbao Wu' 'Deke Guo' 'Chen Xu'] Abstract—Federated learning (FL) is a collaborative machine learning approach that enables multiple clients to train models without sharing their private data. With the rise of deep…
Tingting Chen, Hongming Li, Hao Zheng, Yong Fan
Characterizing brain dynamic functional connectivity (dFC) patterns from functional Magnetic Resonance Imaging (fMRI) data is of paramount importance in imaging neuroscience and medicine. Recently, many graph neural network (GNN) models, combined with transformers or recurrent neural networks (RNNs), have shown great…
Yijingxiu Lu, Sangseon Lee, Soosung Kang, Sun Kim
In recent years, numerous deep learning models have been developed for drug-target interaction (DTI) prediction. These DTI models specialize in handling data with distinct distributions and features, often yielding inconsistent predictions when applied to unseen data points. This inconsistency poses a challenge for…
Ryotaro Kawata, Kohsei Matsutani, Yuri Kinoshita, Naoki Nishikawa + 1 more
Mixture of Experts (MoE), an ensemble of specialized models equipped with a router that dynamically distributes each input to appropriate experts, has achieved successful results in the field of machine learning. However, theoretical understanding of this architecture is falling behind due to its inherent complexity.…
Yuhao Liu, Marzieh Ajirak, Petar M. Djurić
—In this paper, we propose novel Gaussian process-gated hierarchical mixtures of experts (GPHMEs) that are used for building gates and experts. Unlike in other mixtures of experts where the gating models are linear to the input, the gating functions of our model are inner nodes built with Gaussian processes based on…
Bruce Rushing
Construction Authors: ['Bruce Rushing'] Mixture of experts is a prediction aggregation method in machine learning that aggregates the predictions of specialized experts. This method often outperforms Bayesian methods despite the Bayesian having stronger inductive guarantees. We argue that this is due to the greater…
Farhad Zamani, Asta Mannstaedt Rasmussen, Viktoria Schuster, Mathilde Hartvig Diekema + 2 more
MicroRNAs (miRNAs) are important post-transcriptional regulators, yet their expression is typically unobserved in single-cell and most bulk RNA-seq datasets. We present miDGD, a deep generative decoder model that predicts miRNA abundance directly from gene expression alone. Trained on bulk and single-cell datasets from…
Linxing Preston Jiang, Shirui Chen, Emmanuel Tanumihardja, Xiaochuang Han + 3 more
A key challenge in analyzing neuroscience datasets is the profound variability they exhibit across sessions, animals, and data modalities–i.e., heterogeneity. Several recent studies have demonstrated performance gains from pretraining neural foundation models on multi-session datasets, seemingly overcoming this…