20 papers · ranked by Valyu relevance
Zhao, Guoliang, Fu, Yuhan + 22 more
Mixture-of-Experts (MoE) models have become the consensus approach for enabling parameter-efficient scaling and cost-effective deployment in large language models. However, existing scaling laws for dense models are inapplicable to MoE models, which stems from three critical challenges: the multiplicity of influencing…
Jan Ludziejewski, Maciej Pióro, Jakub Krajewski, M. Stefaniak + 7 more
'Michał Krutul' 'Jan Małaśnicki' 'Marek Cygan' 'Piotr Sankowski' 'Kamil Adamczewski' 'Piotr Miłoś' 'Sebastian Jaszczur'] Mixture of Experts (MoE) architectures have significantly increased computational efficiency in both research and real-world applications of large-scale machine learning models. However, their…
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski, Maciej Pióro + 8 more
'Michał Krutul' 'Szymon Antoniak' 'Kamil Ciebiera' 'Krystian Król' 'Tomasz Odrzygóźdź' 'Piotr Sankowski' 'Marek Cygan' 'Sebastian Jaszczur'] | Jakub Krajewski ∗ | Jan Ludziejewski ∗ | Kamil Adamczewski | Maciej Pioro ´ | | --- | --- | --- | --- | | University of Warsaw | University of Warsaw | IDEAS NCBR | IPPT PAN | |…
Youngseog Chung, Dhruv Malik, Jeff Schneider, Yuanzhi Li + 1 more
'Aarti Singh'] The traditional viewpoint on Sparse Mixture of Experts (MoE) models is that instead of training a single large expert, which is computationally expensive, we can train many small experts. The hope is that if the total parameter count of the small experts equals that of the singular large expert, then we…
Ning Sun, Shuxian Zou, Tianhua Tao, Sazan Mahbub + 6 more
Proteins play a fundamental role in life. Understanding the language of proteins offers significant potential for gaining mechanistic insights into biological systems and introduces new avenues for treating diseases, enhancing agriculture, and safeguarding the environment. While large protein language models (PLMs)…
Zhan, Zheng, Ren, Liliang + 12 more
Linear State Space Models (SSMs) offer remarkable performance gains in efficient sequence modeling, with constant inference-time computation and memory complexity. Recent advances, such as Mamba, further enhance SSMs with inputdependent gating and hardware-aware implementations, positioning them as strong alternatives…
Shadi Zabad, Yue Li, Simon Gravel
With the increasing availability of high quality genomic data from diverse cohorts, polygenic scores (PRS) have become a mainstay of genetic analyses of complex traits and diseases. Despite their proliferation in numerous research domains, a major obstacle to wider adoption in clinical settings has been the…
Jialin Wen, Xiaojun Li, Junping Yao, Xinyan Kong + 1 more
Load imbalance is a major performance bottleneck in training mixture-of-experts (MoE) models, as unbalanced expert loads can lead to routing collapse. Most existing approaches address this issue by introducing auxiliary loss functions to balance the load; however, the hyperparameters within these loss functions often…
Xiang Zhang, Shenbao Yu, Jie Xia, Fan Yang
Recent advancements in large-scale self-supervised pretraining have significantly improved molecular representation learning, yet challenges persist, particularly when addressing distributional shifts (e.g., under scaffold-split). Drawing inspiration from the success of Mixture-of-Experts (MoE) networks in NLP, we…
Antonios Vogiatzis, Stavros Orfanoudakis, Georgios Chalkiadakis, Konstantia Moirogiorgou + 2 more
'Konstantia Moirogiorgou' 'Michalis Zervakis' 'Loris Nanni'] Multiclass image classification is a complex task that has been thoroughly investigated in the past. Decomposition-based strategies are commonly employed to address it. Typically, these methods divide the original problem into smaller, potentially simpler…
Jaemoo Hong, Keon Myung Lee, Heming Jia
As recent Multi-Layer Perceptron (MLP) mixer models have achieved state-of-the-art performance in time series forecasting, modeling each MLP-mixer as a separate expert within a mixture is expected to extend the representational capacity of the model, allowing each expert to be activated in response to time-varying…
Authors not listed
Meta-GGA density functional theory (DFT) is an important method in ab initio materials modelling; however, its computational cost limits applicability for generating large datasets or simulating extended length and time scales, as necessary for modern materials discovery. Deorbitalization is a promising strategy to…
Billy Peralta, Ariel Saavedra, Luis Caro, Alvaro Soto
Today, there is growing interest in the automatic classification of a variety of tasks, such as weather forecasting, product recommendations, intrusion detection, and people recognition. “Mixture-of-experts” is a well-known classification technique; it is a probabilistic model consisting of local expert classifiers…
Tingting Chen, Hongming Li, Hao Zheng, Yong Fan
Characterizing brain dynamic functional connectivity (dFC) patterns from functional Magnetic Resonance Imaging (fMRI) data is of paramount importance in imaging neuroscience and medicine. Recently, many graph neural network (GNN) models, combined with transformers or recurrent neural networks (RNNs), have shown great…
Yijingxiu Lu, Sangseon Lee, Soosung Kang, Sun Kim
In recent years, numerous deep learning models have been developed for drug-target interaction (DTI) prediction. These DTI models specialize in handling data with distinct distributions and features, often yielding inconsistent predictions when applied to unseen data points. This inconsistency poses a challenge for…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Authors not listed
Computational toxicology plays a pivotal role in modern drug discovery and environmental risk assessment; however, the reliability of predictive models on unseen chemical scaffolds remains a critical bottleneck. Deep learning architectures, despite their prevalence, are susceptible to ’silent failures’—yielding…
Farhad Zamani, Asta Mannstaedt Rasmussen, Viktoria Schuster, Mathilde Hartvig Diekema + 2 more
MicroRNAs (miRNAs) are important post-transcriptional regulators, yet their expression is typically unobserved in single-cell and most bulk RNA-seq datasets. We present miDGD, a deep generative decoder model that predicts miRNA abundance directly from gene expression alone. Trained on bulk and single-cell datasets from…
Nathan Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi + 3 more
Massive scale, both in terms of data availability and computation, enables significant breakthroughs in key application areas of deep learning such as natural language processing (NLP) and computer vision. There is emerging evidence that scale may be a key ingredient in scientific deep learning, but the importance of…
Authors not listed
This study presents a validation and refinement of the “yellow cards” error detection workflow that can be applied to any property connected to molecular structure. In our implementation the workflow employed 5 predictive models with each assigning a “yellow card” to 5% of the entries with worst prediction accuracy.…