15 papers · ranked by Valyu relevance
Chaoyang He, Shuai Zheng, Aston Zhang, George Karypis + 3 more
'Trishul Chilimbi' 'Mahdi Soltanolkotabi' 'Salman Avestimehr'] The mixture of Expert (MoE) parallelism is a recent advancement that scales up the model size with constant computational cost. MoE selects different sets of parameters (i.e., experts) for each incoming token, resulting in a sparsely-activated model.…
Zhao, Guoliang, Fu, Yuhan + 22 more
Mixture-of-Experts (MoE) models have become the consensus approach for enabling parameter-efficient scaling and cost-effective deployment in large language models. However, existing scaling laws for dense models are inapplicable to MoE models, which stems from three critical challenges: the multiplicity of influencing…
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski, Maciej Pióro + 8 more
'Michał Krutul' 'Szymon Antoniak' 'Kamil Ciebiera' 'Krystian Król' 'Tomasz Odrzygóźdź' 'Piotr Sankowski' 'Marek Cygan' 'Sebastian Jaszczur'] | Jakub Krajewski ∗ | Jan Ludziejewski ∗ | Kamil Adamczewski | Maciej Pioro ´ | | --- | --- | --- | --- | | University of Warsaw | University of Warsaw | IDEAS NCBR | IPPT PAN | |…
Youngseog Chung, Dhruv Malik, Jeff Schneider, Yuanzhi Li + 1 more
'Aarti Singh'] The traditional viewpoint on Sparse Mixture of Experts (MoE) models is that instead of training a single large expert, which is computationally expensive, we can train many small experts. The hope is that if the total parameter count of the small experts equals that of the singular large expert, then we…
Lemeng Wu, Mengchen Liu, Yinpeng Chen, Dongdong Chen + 2 more
'Lu Yuan'] Mixture of Experts (MoE) is able to scale up vision transformers effectively. However, it requires prohibiting computation resources to train a large MoE transformer. In this paper, we propose Residual Mixture of Experts (RMoE), an efficient training pipeline for MoE vision transformers on downstream tasks…
James Oldfield, Markos Georgopoulos, Grigorios G. Chrysos, Christos Tzelepis + 4 more
Factorization Authors: ['James Oldfield' 'Markos Georgopoulos' 'Grigorios G. Chrysos' 'Christos Tzelepis' 'Yannis Panagakis' 'Mihalis A. Nicolaou' 'Jiankang Deng' 'Ioannis Patras'] The Mixture of Experts (MoE) paradigm provides a powerful way to decompose inscrutable dense layers into smaller, modular computations…
Songhao Wu, Ang Lv, Ruobing Xie, Yankai Lin
Router is the cornerstone component to the Mixture-of-Experts models. Serving as expert proxies, the rows of the router matrix compute their similarity to the MoE inputs to determine which subset of experts is activated. Ideally, each router row is designed to encode the expert matrix into this representative vector…
Runxi Cheng, Yuchen Guan, Yucheng Ding, Qingguo Hu + 5 more
In this work, We first explore whether the parameters activated by the MoE layer remain highly sparse at inference. We perform a sparsification study on several representative MoE models. For each expert, we rank parameters by the magnitude of their activations from the gate projection and progressively prune the…
Zefeng Gao, Peiyu Liu, Wayne Xin Zhao, LU Zhong-yi + 1 more
Recently, Mixture-of-Experts (short as MoE) architecture has achieved remarkable success in increasing the model capacity of largescale language models. However, MoE requires incorporating significantly more parameters than the base model being extended. In this paper, we propose building a parameterefficient MoE…
Hongbo Li, Qiqi Wu, Sen Lin, Yingbin Liang + 1 more
Mixture-of-Experts (MoE) models improve transformer efficiency but lack a unified theoretical explanation—especially when both feed-forward and attention layers are allowed to specialize. To this end, we study the Mixture-of-Transformers (MoT), a tractable theoretical framework in which each transformer block acts as…
Ziwei Zhan, Wei Zhao, Yuanqing Li, Weijie Liu + 5 more
Aggregation Authors: ['Ziwei Zhan' 'Wei Zhao' 'Yuanqing Li' 'Weijie Liu' 'X.C. Zhang' 'Chee Wei Tan' 'Chuanbao Wu' 'Deke Guo' 'Chen Xu'] Abstract—Federated learning (FL) is a collaborative machine learning approach that enables multiple clients to train models without sharing their private data. With the rise of deep…
Jessica Leoni, Valentina Breschi, Simone Formentin, Mara Tanelli
effective blending of grey and black-box models Authors: ['Jessica Leoni' 'Valentina Breschi' 'Simone Formentin' 'Mara Tanelli'] Traditional models grounded in first principles often struggle with accuracy as the system's complexity increases. Conversely, machine learning approaches, while powerful, face challenges in…
Ryotaro Kawata, Kohsei Matsutani, Yuri Kinoshita, Naoki Nishikawa + 1 more
Mixture of Experts (MoE), an ensemble of specialized models equipped with a router that dynamically distributes each input to appropriate experts, has achieved successful results in the field of machine learning. However, theoretical understanding of this architecture is falling behind due to its inherent complexity.…
Yuhao Liu, Marzieh Ajirak, Petar M. Djurić
—In this paper, we propose novel Gaussian process-gated hierarchical mixtures of experts (GPHMEs) that are used for building gates and experts. Unlike in other mixtures of experts where the gating models are linear to the input, the gating functions of our model are inner nodes built with Gaussian processes based on…
Bruce Rushing
Construction Authors: ['Bruce Rushing'] Mixture of experts is a prediction aggregation method in machine learning that aggregates the predictions of specialized experts. This method often outperforms Bayesian methods despite the Bayesian having stronger inductive guarantees. We argue that this is due to the greater…