25 papers · ranked by Valyu relevance
Antonios Vogiatzis, Stavros Orfanoudakis, Georgios Chalkiadakis, Konstantia Moirogiorgou + 2 more
'Konstantia Moirogiorgou' 'Michalis Zervakis' 'Loris Nanni'] Multiclass image classification is a complex task that has been thoroughly investigated in the past. Decomposition-based strategies are commonly employed to address it. Typically, these methods divide the original problem into smaller, potentially simpler…
Billy Peralta, Ariel Saavedra, Luis Caro, Alvaro Soto
Today, there is growing interest in the automatic classification of a variety of tasks, such as weather forecasting, product recommendations, intrusion detection, and people recognition. “Mixture-of-experts” is a well-known classification technique; it is a probabilistic model consisting of local expert classifiers…
Shadi Zabad, Yue Li, Simon Gravel
With the increasing availability of high quality genomic data from diverse cohorts, polygenic scores (PRS) have become a mainstay of genetic analyses of complex traits and diseases. Despite their proliferation in numerous research domains, a major obstacle to wider adoption in clinical settings has been the…
Runxi Cheng, Yuchen Guan, Yucheng Ding, Qingguo Hu + 5 more
In this work, We first explore whether the parameters activated by the MoE layer remain highly sparse at inference. We perform a sparsification study on several representative MoE models. For each expert, we rank parameters by the magnitude of their activations from the gate projection and progressively prune the…
Yijingxiu Lu, Sangseon Lee, Soosung Kang, Sun Kim
In recent years, numerous deep learning models have been developed for drug-target interaction (DTI) prediction. These DTI models specialize in handling data with distinct distributions and features, often yielding inconsistent predictions when applied to unseen data points. This inconsistency poses a challenge for…
Billy Peralta
A useful strategy to deal with complex classification scenarios is the "divide and conquer" approach. The mixture of experts (MOE) technique makes use of this strategy by joinly training a set of classifiers, or experts, that are specialized in different regions of the input space. A global model, or gate function…
Yuxi Liu, Zhenhao Zhang, Mufan Qiu, Song Wang + 5 more
Single-cell RNA sequencing (scRNA-seq) enables high-resolution characterization of cellular heterogeneity, but its rich, complementary structure across cells and genes remains underexploited, especially in the presence of technical noise and sparsity. Effectively leveraging this multi-scale structure is essentially an…
Yanjun Qi, Judith Klein-Seetharaman, Ziv Bar-Joseph
Background High-throughput methods can directly detect the set of interacting proteins in model species but the results are often incomplete and exhibit high false positive and false negative rates. A number of researchers have recently presented methods for integrating direct and indirect data for predicting…
Jialin Wen, Xiaojun Li, Junping Yao, Xinyan Kong + 1 more
Load imbalance is a major performance bottleneck in training mixture-of-experts (MoE) models, as unbalanced expert loads can lead to routing collapse. Most existing approaches address this issue by introducing auxiliary loss functions to balance the load; however, the hyperparameters within these loss functions often…
Manxi Sun, Wei Liu, Jing Luan, Pengfei Gao + 1 more
The Sparsely-Activated Mixture-of-Experts (MoE) has gained increasing popularity for scaling up large language models (LLMs) without exploding computational costs. Despite its success, the current design faces a challenge where all experts have the same size, limiting the ability of tokens to choose the experts with…
Sudhir Raman, Thomas J Fuchs, Peter J Wild, Edgar Dahl + 2 more
Background We present an infinite mixture-of-experts model to find an unknown number of sub-groups within a given patient cohort based on survival analysis. The effect of patient features on survival is modeled using the Cox’s proportionality hazards model which yields a non-standard regression component. The model is…
Yizhak Ben-Shabat, Chamin Hewa Koneputugodage, Sameera Ramasinghe, Stephen Jay Gould
'Stephen Jay Gould'] Implicit neural representations (INRs) have proven effective in various tasks including image, shape, audio, and video reconstruction. These INRs typically learn the implicit field from sampled input points. This is often done using a single network for the entire domain, imposing many global…
Jaemoo Hong, Keon Myung Lee, Heming Jia
As recent Multi-Layer Perceptron (MLP) mixer models have achieved state-of-the-art performance in time series forecasting, modeling each MLP-mixer as a separate expert within a mixture is expected to extend the representational capacity of the model, allowing each expert to be activated in response to time-varying…
Bruce Rushing
Construction Authors: ['Bruce Rushing'] Mixture of experts is a prediction aggregation method in machine learning that aggregates the predictions of specialized experts. This method often outperforms Bayesian methods despite the Bayesian having stronger inductive guarantees. We argue that this is due to the greater…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Marie Courbariaux, Kylliann De Santiago, Cyril Dalmasso, Fabrice Danjou + 5 more
'Fabrice Danjou' 'Samir Bekadar' 'Jean-Christophe Corvol' 'Maria Martinez' 'Marie Szafranski' 'Christophe Ambroise'] Motivation: Identifying new genetic associations in non-Mendelian complex diseases is an increasingly difficult challenge. These diseases sometimes appear to have a significant component of heritability…
Authors not listed
Meta-GGA density functional theory (DFT) is an important method in ab initio materials modelling; however, its computational cost limits applicability for generating large datasets or simulating extended length and time scales, as necessary for modern materials discovery. Deorbitalization is a promising strategy to…
Xiang Zhang, Shenbao Yu, Jie Xia, Fan Yang
Recent advancements in large-scale self-supervised pretraining have significantly improved molecular representation learning, yet challenges persist, particularly when addressing distributional shifts (e.g., under scaffold-split). Drawing inspiration from the success of Mixture-of-Experts (MoE) networks in NLP, we…
Tingting Chen, Hongming Li, Hao Zheng, Yong Fan
Characterizing brain dynamic functional connectivity (dFC) patterns from functional Magnetic Resonance Imaging (fMRI) data is of paramount importance in imaging neuroscience and medicine. Recently, many graph neural network (GNN) models, combined with transformers or recurrent neural networks (RNNs), have shown great…
Wenbo Zhao, Yang Gao, Shahan Ali Memon, Bhiksha Raj + 1 more
In regression tasks the distribution of the data is often too complex to be fitted by a single model. In contrast, partition-based models are developed where data is divided and fitted by local models. These models partition the input space and do not leverage the input-output dependency of multimodal-distributed data…
Yanshuai Cao, David J. Fleet
In this work, we propose a generalized product of experts (gPoE) framework for combining the predictions of multiple probabilistic models. We identify four desirable properties that are important for scalability, expressiveness and robustness, when learning and inferring with a combination of multiple models. Through…
Authors not listed
Computational toxicology plays a pivotal role in modern drug discovery and environmental risk assessment; however, the reliability of predictive models on unseen chemical scaffolds remains a critical bottleneck. Deep learning architectures, despite their prevalence, are susceptible to ’silent failures’—yielding…
Oscar Oelrich, Mattias Villani, Sebastian Ankargren
We propose local prediction pools as a method for combining the predictive distributions of a set of experts conditional on a set of variables believed to be related to the predictive accuracy of the experts. This is done in a two step process where we first estimate the conditional predictive accuracy of each expert…
Authors not listed
The accurate prediction of fuel mixture properties is essential for the development of alternative fuels, yet remains challenging under data-scarce conditions due to the combinatorial complexity of multi-component systems. In this study, we present a systematic evaluation of three machine learning (ML)…
Yasmine Nahal, Janosch Menke, Julien Martinelli, Markus Heinonen + 5 more
Machine learning (ML) systems have enabled the modelling of quantitative structure-property relationships (QSPR) and structure-activity relationships (QSAR) using existing experimental data to predict target properties for new molecules. These property predictors hold significant potential in accelerating drug…