23 papers · ranked by Valyu relevance
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski, Maciej Pióro + 8 more
'Michał Krutul' 'Szymon Antoniak' 'Kamil Ciebiera' 'Krystian Król' 'Tomasz Odrzygóźdź' 'Piotr Sankowski' 'Marek Cygan' 'Sebastian Jaszczur'] | Jakub Krajewski ∗ | Jan Ludziejewski ∗ | Kamil Adamczewski | Maciej Pioro ´ | | --- | --- | --- | --- | | University of Warsaw | University of Warsaw | IDEAS NCBR | IPPT PAN | |…
Youngseog Chung, Dhruv Malik, Jeff Schneider, Yuanzhi Li + 1 more
'Aarti Singh'] The traditional viewpoint on Sparse Mixture of Experts (MoE) models is that instead of training a single large expert, which is computationally expensive, we can train many small experts. The hope is that if the total parameter count of the small experts equals that of the singular large expert, then we…
Ning Sun, Shuxian Zou, Tianhua Tao, Sazan Mahbub + 6 more
Proteins play a fundamental role in life. Understanding the language of proteins offers significant potential for gaining mechanistic insights into biological systems and introduces new avenues for treating diseases, enhancing agriculture, and safeguarding the environment. While large protein language models (PLMs)…
James Oldfield, Markos Georgopoulos, Grigorios G. Chrysos, Christos Tzelepis + 4 more
Factorization Authors: ['James Oldfield' 'Markos Georgopoulos' 'Grigorios G. Chrysos' 'Christos Tzelepis' 'Yannis Panagakis' 'Mihalis A. Nicolaou' 'Jiankang Deng' 'Ioannis Patras'] The Mixture of Experts (MoE) paradigm provides a powerful way to decompose inscrutable dense layers into smaller, modular computations…
Shadi Zabad, Yue Li, Simon Gravel
With the increasing availability of high quality genomic data from diverse cohorts, polygenic scores (PRS) have become a mainstay of genetic analyses of complex traits and diseases. Despite their proliferation in numerous research domains, a major obstacle to wider adoption in clinical settings has been the…
Xiang Zhang, Shenbao Yu, Jie Xia, Fan Yang
Recent advancements in large-scale self-supervised pretraining have significantly improved molecular representation learning, yet challenges persist, particularly when addressing distributional shifts (e.g., under scaffold-split). Drawing inspiration from the success of Mixture-of-Experts (MoE) networks in NLP, we…
Jialin Wen, Xiaojun Li, Junping Yao, Xinyan Kong + 1 more
Load imbalance is a major performance bottleneck in training mixture-of-experts (MoE) models, as unbalanced expert loads can lead to routing collapse. Most existing approaches address this issue by introducing auxiliary loss functions to balance the load; however, the hyperparameters within these loss functions often…
Jaemoo Hong, Keon Myung Lee, Heming Jia
As recent Multi-Layer Perceptron (MLP) mixer models have achieved state-of-the-art performance in time series forecasting, modeling each MLP-mixer as a separate expert within a mixture is expected to extend the representational capacity of the model, allowing each expert to be activated in response to time-varying…
Antonios Vogiatzis, Stavros Orfanoudakis, Georgios Chalkiadakis, Konstantia Moirogiorgou + 2 more
'Konstantia Moirogiorgou' 'Michalis Zervakis' 'Loris Nanni'] Multiclass image classification is a complex task that has been thoroughly investigated in the past. Decomposition-based strategies are commonly employed to address it. Typically, these methods divide the original problem into smaller, potentially simpler…
Ziwei Zhan, Wei Zhao, Yuanqing Li, Weijie Liu + 5 more
Aggregation Authors: ['Ziwei Zhan' 'Wei Zhao' 'Yuanqing Li' 'Weijie Liu' 'X.C. Zhang' 'Chee Wei Tan' 'Chuanbao Wu' 'Deke Guo' 'Chen Xu'] Abstract—Federated learning (FL) is a collaborative machine learning approach that enables multiple clients to train models without sharing their private data. With the rise of deep…
Authors not listed
Meta-GGA density functional theory (DFT) is an important method in ab initio materials modelling; however, its computational cost limits applicability for generating large datasets or simulating extended length and time scales, as necessary for modern materials discovery. Deorbitalization is a promising strategy to…
Tingting Chen, Hongming Li, Hao Zheng, Yong Fan
Characterizing brain dynamic functional connectivity (dFC) patterns from functional Magnetic Resonance Imaging (fMRI) data is of paramount importance in imaging neuroscience and medicine. Recently, many graph neural network (GNN) models, combined with transformers or recurrent neural networks (RNNs), have shown great…
Yijingxiu Lu, Sangseon Lee, Soosung Kang, Sun Kim
In recent years, numerous deep learning models have been developed for drug-target interaction (DTI) prediction. These DTI models specialize in handling data with distinct distributions and features, often yielding inconsistent predictions when applied to unseen data points. This inconsistency poses a challenge for…
Ryotaro Kawata, Kohsei Matsutani, Yuri Kinoshita, Naoki Nishikawa + 1 more
Mixture of Experts (MoE), an ensemble of specialized models equipped with a router that dynamically distributes each input to appropriate experts, has achieved successful results in the field of machine learning. However, theoretical understanding of this architecture is falling behind due to its inherent complexity.…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Yuhao Liu, Marzieh Ajirak, Petar M. Djurić
—In this paper, we propose novel Gaussian process-gated hierarchical mixtures of experts (GPHMEs) that are used for building gates and experts. Unlike in other mixtures of experts where the gating models are linear to the input, the gating functions of our model are inner nodes built with Gaussian processes based on…
Bruce Rushing
Construction Authors: ['Bruce Rushing'] Mixture of experts is a prediction aggregation method in machine learning that aggregates the predictions of specialized experts. This method often outperforms Bayesian methods despite the Bayesian having stronger inductive guarantees. We argue that this is due to the greater…
Authors not listed
Computational toxicology plays a pivotal role in modern drug discovery and environmental risk assessment; however, the reliability of predictive models on unseen chemical scaffolds remains a critical bottleneck. Deep learning architectures, despite their prevalence, are susceptible to ’silent failures’—yielding…
Yingshuai Wang, Dezheng Zhang, Aziguli Wulamu
Training models to predict click and order targets at the same time. For better user satisfaction and business effectiveness, multitask learning is one of the most important methods in e-commerce. Some existing researches model user representation based on historical behaviour sequence to capture user interests. It is…
Hamed Jalali, Gjergji Kasneci, Robertas Alzbutas, Mark Girolami + 1 more
'Hussein Rappel'] By distributing the training process, local approximation reduces the cost of the standard Gaussian process. An ensemble method aggregates predictions from local Gaussian experts, each trained on different data partitions, under the assumption of perfect diversity among them. While this assumption…
Nathan Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi + 3 more
Massive scale, both in terms of data availability and computation, enables significant breakthroughs in key application areas of deep learning such as natural language processing (NLP) and computer vision. There is emerging evidence that scale may be a key ingredient in scientific deep learning, but the importance of…
Farhad Zamani, Asta Mannstaedt Rasmussen, Viktoria Schuster, Mathilde Hartvig Diekema + 2 more
MicroRNAs (miRNAs) are important post-transcriptional regulators, yet their expression is typically unobserved in single-cell and most bulk RNA-seq datasets. We present miDGD, a deep generative decoder model that predicts miRNA abundance directly from gene expression alone. Trained on bulk and single-cell datasets from…
Linxing Preston Jiang, Shirui Chen, Emmanuel Tanumihardja, Xiaochuang Han + 3 more
A key challenge in analyzing neuroscience datasets is the profound variability they exhibit across sessions, animals, and data modalities–i.e., heterogeneity. Several recent studies have demonstrated performance gains from pretraining neural foundation models on multi-session datasets, seemingly overcoming this…