11 papers · ranked by Valyu relevance
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski, Maciej Pióro + 8 more
'Michał Krutul' 'Szymon Antoniak' 'Kamil Ciebiera' 'Krystian Król' 'Tomasz Odrzygóźdź' 'Piotr Sankowski' 'Marek Cygan' 'Sebastian Jaszczur'] | Jakub Krajewski ∗ | Jan Ludziejewski ∗ | Kamil Adamczewski | Maciej Pioro ´ | | --- | --- | --- | --- | | University of Warsaw | University of Warsaw | IDEAS NCBR | IPPT PAN | |…
Youngseog Chung, Dhruv Malik, Jeff Schneider, Yuanzhi Li + 1 more
'Aarti Singh'] The traditional viewpoint on Sparse Mixture of Experts (MoE) models is that instead of training a single large expert, which is computationally expensive, we can train many small experts. The hope is that if the total parameter count of the small experts equals that of the singular large expert, then we…
James Oldfield, Markos Georgopoulos, Grigorios G. Chrysos, Christos Tzelepis + 4 more
Factorization Authors: ['James Oldfield' 'Markos Georgopoulos' 'Grigorios G. Chrysos' 'Christos Tzelepis' 'Yannis Panagakis' 'Mihalis A. Nicolaou' 'Jiankang Deng' 'Ioannis Patras'] The Mixture of Experts (MoE) paradigm provides a powerful way to decompose inscrutable dense layers into smaller, modular computations…
Ziwei Zhan, Wei Zhao, Yuanqing Li, Weijie Liu + 5 more
Aggregation Authors: ['Ziwei Zhan' 'Wei Zhao' 'Yuanqing Li' 'Weijie Liu' 'X.C. Zhang' 'Chee Wei Tan' 'Chuanbao Wu' 'Deke Guo' 'Chen Xu'] Abstract—Federated learning (FL) is a collaborative machine learning approach that enables multiple clients to train models without sharing their private data. With the rise of deep…
Authors not listed
Meta-GGA density functional theory (DFT) is an important method in ab initio materials modelling; however, its computational cost limits applicability for generating large datasets or simulating extended length and time scales, as necessary for modern materials discovery. Deorbitalization is a promising strategy to…
Ryotaro Kawata, Kohsei Matsutani, Yuri Kinoshita, Naoki Nishikawa + 1 more
Mixture of Experts (MoE), an ensemble of specialized models equipped with a router that dynamically distributes each input to appropriate experts, has achieved successful results in the field of machine learning. However, theoretical understanding of this architecture is falling behind due to its inherent complexity.…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Yuhao Liu, Marzieh Ajirak, Petar M. Djurić
—In this paper, we propose novel Gaussian process-gated hierarchical mixtures of experts (GPHMEs) that are used for building gates and experts. Unlike in other mixtures of experts where the gating models are linear to the input, the gating functions of our model are inner nodes built with Gaussian processes based on…
Bruce Rushing
Construction Authors: ['Bruce Rushing'] Mixture of experts is a prediction aggregation method in machine learning that aggregates the predictions of specialized experts. This method often outperforms Bayesian methods despite the Bayesian having stronger inductive guarantees. We argue that this is due to the greater…
Authors not listed
Computational toxicology plays a pivotal role in modern drug discovery and environmental risk assessment; however, the reliability of predictive models on unseen chemical scaffolds remains a critical bottleneck. Deep learning architectures, despite their prevalence, are susceptible to ’silent failures’—yielding…
Nathan Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi + 3 more
Massive scale, both in terms of data availability and computation, enables significant breakthroughs in key application areas of deep learning such as natural language processing (NLP) and computer vision. There is emerging evidence that scale may be a key ingredient in scientific deep learning, but the importance of…