25 papers · ranked by Valyu relevance
Zeng, Zhichen, Hang, Mengyue + 24 more
Deep models have driven significant advances in click-through rate (CTR) prediction. While vertical scaling via layer stacking improves model expressiveness, the layer-by-layer sequential computation poses challenges to efficient scaling. Conversely, horizontal scaling through Mixture of Experts (MoE) achieves…
Leena Chennuru Vankadara, Moritz Haas, Luke Hayward, Sebastian Bordt + 1 more
Recent frontier large language models predominantly rely on Mixture-of-Experts (MoE) architectures. Despite empirical progress, there is still no principled understanding of how hyperparameters should scale with network width $N$, expert width $N_e$, number of experts $M$, sparsity $K$, and depth $L$ to ensure both…
Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang + 23 more
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visual fidelity and creativity over computational efficiency and physical realism. In this work, we present LingBot-Video, a…
Qi Zhou, Yuan Le, Xiaoning Qi, Hanshuo Lu + 5 more
Foundation models learned from single-cell transcriptomes are central to the prospect of AI virtual cell that can represent, query and predict cellular state. However, most current single-cell foundation models learn from a single view of gene expression and are optimized primarily through reconstruction or next-token…
Dong Sun, Rahul Nittala, Rebekka Burkholz
Despite their practical success, it remains unclear why Mixture of Experts (MoE) models can outperform dense networks beyond sheer parameter scaling. We study an iso-parameter regime where inputs exhibit latent modular structure but are corrupted by feature noise, a proxy for noisy internal activations. We show that…
Jialin Wen, Xiaojun Li, Junping Yao, Xinyan Kong + 1 more
Load imbalance is a major performance bottleneck in training mixture-of-experts (MoE) models, as unbalanced expert loads can lead to routing collapse. Most existing approaches address this issue by introducing auxiliary loss functions to balance the load; however, the hyperparameters within these loss functions often…
Feilong Liu
Mixture-of-Experts (MoE) architectures are commonly motivated by efficiency and conditional computation, but their effect on the geometry of learned functions and representations remains poorly characterized. In this work, we study MoEs through a geometric lens, interpreting routing as a form of soft partitioning of…
Runxi Cheng, Yuchen Guan, Yucheng Ding, Qingguo Hu + 5 more
In this work, We first explore whether the parameters activated by the MoE layer remain highly sparse at inference. We perform a sparsification study on several representative MoE models. For each expert, we rank parameters by the magnitude of their activations from the gate projection and progressively prune the…
Sultan Sevgi Turgut Ögme, Nizamettin Aydin, Zeyneb Kurt, Hatice Ulku Osmanbeyoglu
Single-cell RNA-seq (scRNAseq) analyses performed at the cellular level aim to understand the cellular landscape of tissue sections, offer insights into rare cell-types, and identify marker genes for annotating distinct cell types. ScRNAseq analyses are widely applied to cancer research to understand tumor…
Songhao Wu, Ang Lv, Ruobing Xie, Yankai Lin
Router is the cornerstone component to the Mixture-of-Experts models. Serving as expert proxies, the rows of the router matrix compute their similarity to the MoE inputs to determine which subset of experts is activated. Ideally, each router row is designed to encode the expert matrix into this representative vector…
Authors not listed
Meta-GGA density functional theory (DFT) is an important method in ab initio materials modelling; however, its computational cost limits applicability for generating large datasets or simulating extended length and time scales, as necessary for modern materials discovery. Deorbitalization is a promising strategy to…
Dipayan Sarkar, Chiranjib Sarkar
De novo protein binder design has been dominated by structure-based pipelines that require known three-dimensional target conformations and consume substantial compute and generation time per design, limiting their throughput and accessibility for routine large-scale binder exploration. Sequence-only generative models…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Shradhanjali Das, K. Hemalatha
Medical image analysis faces persistent challenges due to the distributed data, limited annotations, and variations in imaging modalities, acquisition protocols, and patient demographics. Centralized deep learning approaches compromise data privacy, while Federated Learning (FL) enables decentralized model training…
Jaemoo Hong, Keon Myung Lee, Heming Jia
As recent Multi-Layer Perceptron (MLP) mixer models have achieved state-of-the-art performance in time series forecasting, modeling each MLP-mixer as a separate expert within a mixture is expected to extend the representational capacity of the model, allowing each expert to be activated in response to time-varying…
Hongyu Xiong, Ming Dai
Existing multimodal federated learning methods typically assume complete modality availability and struggle with heterogeneity between training and testing data distributions, making them unsuitable for handling missing modalities and distribution drift in distributed learning scenarios such as the Internet of Things…
Liang Wang
How do multi-modal large language models that jointly process natural language and biological sequences (DNA, protein, structural alphabets) actually answer biological questions, especially sequence-grounded questions whose answer depends on residue-level patterns rather than literature recall? We introduce OmniGene-4…
Hao Li, Yuyang Feng, Xin Zhao, Xuan Li + 1 more
Person re-identification (re-ID) aims to match pedestrian images across disjoint camera views. Existing multi-source unsupervised domain adaptation (UDA) re-ID methods still face two critical issues: they fail to effectively balance domain-invariant feature learning and domain-specific style preservation and cannot…
Authors not listed
Computational toxicology plays a pivotal role in modern drug discovery and environmental risk assessment; however, the reliability of predictive models on unseen chemical scaffolds remains a critical bottleneck. Deep learning architectures, despite their prevalence, are susceptible to ’silent failures’—yielding…
Authors not listed
Accurate extrapolation in data-scarce scientific systems remains a central challenge for machine intelligence. In microbial bioprocessing, kinetic parameters change non-monotonically with reactor volume due to interacting hydrodynamic, oxygen-transfer, and mixing effects, rendering classical empirical scaling laws…
Aidan Dempster, Brokoslaw Laschowski
A grand challenge in brain decoding is to develop algorithms that generalize across multiple subjects and tasks. Here, we developed a new computational framework to minimize negative transfer for domain-adaptive brain decoding by reframing source selection as a mixture model parameter estimation problem, allowing each…
Francesco G. Rinaldi, Eugenio Piasini
To make sense of a noisy world, living beings constantly face decisions between competing interpretations for ambiguous sensory data. This process parallels statistical model selection, where most frameworks, like the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), are based on a…
Farhad Zamani, Asta Mannstaedt Rasmussen, Viktoria Schuster, Mathilde Hartvig Diekema + 2 more
MicroRNAs (miRNAs) are important post-transcriptional regulators, yet their expression is typically unobserved in single-cell and most bulk RNA-seq datasets. We present miDGD, a deep generative decoder model that predicts miRNA abundance directly from gene expression alone. Trained on bulk and single-cell datasets from…
Authors not listed
Chemical language models, such as transformers trained on SMILES strings, are increasingly used in drug design and have faced rapid growth in both model capacity and training dataset size. The impact of this scaling on practical downstream performance remains unclear, however. We systematically evaluate how model size…
Authors not listed
In this work, we present EquiNet, a neural network for predicting vapor–liquid equilibrium (VLE) in novel binary mixtures through direct estimation of activity coefficients and vapor pressures. The model embeds a classic excess-Gibbs free energy formulation, ensuring Gibbs–Duhem consistency on all predicted activity…