23 papers · ranked by Valyu relevance
Francisco das Chagas de Souza, Tim Offermans, Ruud Barendse, Geert Postma + 1 more
'Geert Postma' 'Jeroen Jansen'] Abstract—This work proposes a new data-driven model devised to integrate process knowledge into its structure to increase the human-machine synergy in the process industry. The proposed Contextual Mixture of Experts (cMoE) explicitly uses process knowledge along the model learning stage…
Antonios Vogiatzis, Stavros Orfanoudakis, Georgios Chalkiadakis, Konstantia Moirogiorgou + 2 more
'Konstantia Moirogiorgou' 'Michalis Zervakis' 'Loris Nanni'] Multiclass image classification is a complex task that has been thoroughly investigated in the past. Decomposition-based strategies are commonly employed to address it. Typically, these methods divide the original problem into smaller, potentially simpler…
Shadi Zabad, Yue Li, Simon Gravel
With the increasing availability of high quality genomic data from diverse cohorts, polygenic scores (PRS) have become a mainstay of genetic analyses of complex traits and diseases. Despite their proliferation in numerous research domains, a major obstacle to wider adoption in clinical settings has been the…
Runxi Cheng, Yuchen Guan, Yucheng Ding, Qingguo Hu + 5 more
In this work, We first explore whether the parameters activated by the MoE layer remain highly sparse at inference. We perform a sparsification study on several representative MoE models. For each expert, we rank parameters by the magnitude of their activations from the gate projection and progressively prune the…
Yijingxiu Lu, Sangseon Lee, Soosung Kang, Sun Kim
In recent years, numerous deep learning models have been developed for drug-target interaction (DTI) prediction. These DTI models specialize in handling data with distinct distributions and features, often yielding inconsistent predictions when applied to unseen data points. This inconsistency poses a challenge for…
Yuxi Liu, Zhenhao Zhang, Mufan Qiu, Song Wang + 5 more
Single-cell RNA sequencing (scRNA-seq) enables high-resolution characterization of cellular heterogeneity, but its rich, complementary structure across cells and genes remains underexploited, especially in the presence of technical noise and sparsity. Effectively leveraging this multi-scale structure is essentially an…
Jialin Wen, Xiaojun Li, Junping Yao, Xinyan Kong + 1 more
Load imbalance is a major performance bottleneck in training mixture-of-experts (MoE) models, as unbalanced expert loads can lead to routing collapse. Most existing approaches address this issue by introducing auxiliary loss functions to balance the load; however, the hyperparameters within these loss functions often…
Uthsav Chitra, Shu Dan, Fenna Krienen, Benjamin J. Raphael
Gene expression varies across a tissue due to both the organization of the tissue into spatial domains, i.e. discrete regions of a tissue with distinct cell type composition, and continuous spatial gradients of gene expression within different spatial domains. Spatially resolved transcriptomics (SRT) technologies…
Marie Courbariaux, Kylliann De Santiago, Cyril Dalmasso, Fabrice Danjou + 5 more
'Fabrice Danjou' 'Samir Bekadar' 'Jean-Christophe Corvol' 'Maria Martinez' 'Marie Szafranski' 'Christophe Ambroise'] Motivation: Identifying new genetic associations in non-Mendelian complex diseases is an increasingly difficult challenge. These diseases sometimes appear to have a significant component of heritability…
NATHAN C. HURLEY, SANKET S. DHRUVA, NIHAR R. DESAI, JOSEPH R. ROSS + 4 more
'CHE G. NGUFOR' 'FREDERICK MASOUDI' 'HARLAN M. KRUMHOLZ' 'BOBAK J. MORTAZAVI'] Observational medical data present unique opportunities for analysis of medical outcomes and treatment decision making. However, because these datasets do not contain the strict pairing of randomized control trials, matching techniques are…
Jaemoo Hong, Keon Myung Lee, Heming Jia
As recent Multi-Layer Perceptron (MLP) mixer models have achieved state-of-the-art performance in time series forecasting, modeling each MLP-mixer as a separate expert within a mixture is expected to extend the representational capacity of the model, allowing each expert to be activated in response to time-varying…
Bruce Rushing
Construction Authors: ['Bruce Rushing'] Mixture of experts is a prediction aggregation method in machine learning that aggregates the predictions of specialized experts. This method often outperforms Bayesian methods despite the Bayesian having stronger inductive guarantees. We argue that this is due to the greater…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Authors not listed
Meta-GGA density functional theory (DFT) is an important method in ab initio materials modelling; however, its computational cost limits applicability for generating large datasets or simulating extended length and time scales, as necessary for modern materials discovery. Deorbitalization is a promising strategy to…
Xiang Zhang, Shenbao Yu, Jie Xia, Fan Yang
Recent advancements in large-scale self-supervised pretraining have significantly improved molecular representation learning, yet challenges persist, particularly when addressing distributional shifts (e.g., under scaffold-split). Drawing inspiration from the success of Mixture-of-Experts (MoE) networks in NLP, we…
Forest Mars
Validated domain expertise reliably enhances judgment within its boundaries but creates systematic vulnerabilities at domain borders: experts confronting problems that resemble but causally differ from their training reliably underperform novices facing identical tasks. We term this phenomenon Transitive Expert Error…
Liang Wang
How do multi-modal large language models that jointly process natural language and biological sequences (DNA, protein, structural alphabets) actually answer biological questions, especially sequence-grounded questions whose answer depends on residue-level patterns rather than literature recall? We introduce OmniGene-4…
Jaron T. Colas, John P. O’Doherty, Scott T. Grafton, Stefano Palminteri
Active reinforcement learning enables dynamic prediction and control, where one should not only maximize rewards but also minimize costs such as of inference, decisions, actions, and time. For an embodied agent such as a human, decisions are also shaped by physical aspects of actions. Beyond the effects of reward…
Authors not listed
Computational toxicology plays a pivotal role in modern drug discovery and environmental risk assessment; however, the reliability of predictive models on unseen chemical scaffolds remains a critical bottleneck. Deep learning architectures, despite their prevalence, are susceptible to ’silent failures’—yielding…
Axel Abels, Tom Lenaerts, Vito Trianni, Ann Nowé
Quite some real-world problems can be formulated as decision-making problems wherein one must repeatedly make an appropriate choice from a set of alternatives. Multiple expert judgements, whether human or artificial, can help in taking correct decisions, especially when exploration of alternative solutions is costly.…
Alihan Hüyük, Qiyao Wei, Alicia Curth, Mihaela van der Schaar
Decision-makers are often experts of their domain and take actions based on their domain knowledge. Doctors, for instance, may prescribe treatments by predicting the likely outcome of each available treatment. Actions of an expert thus naturally encode part of their domain knowledge, and can help make inferences within…
Authors not listed
The accurate prediction of fuel mixture properties is essential for the development of alternative fuels, yet remains challenging under data-scarce conditions due to the combinatorial complexity of multi-component systems. In this study, we present a systematic evaluation of three machine learning (ML)…
Axel Abels, Tom Lenaerts, Vito Trianni, Ann Nowé
Experts advising decision-makers are likely to display expertise which varies as a function of the problem instance. In practice, this may lead to sub-optimal or discriminatory decisions against minority cases. In this work we model such changes in depth and breadth of knowledge as a partitioning of the problem space…