14 papers · ranked by Valyu relevance
Xiang Zhang, Shenbao Yu, Jie Xia, Fan Yang
Recent advancements in large-scale self-supervised pretraining have significantly improved molecular representation learning, yet challenges persist, particularly when addressing distributional shifts (e.g., under scaffold-split). Drawing inspiration from the success of Mixture-of-Experts (MoE) networks in NLP, we…
Ning Sun, Shuxian Zou, Tianhua Tao, Sazan Mahbub + 6 more
Proteins play a fundamental role in life. Understanding the language of proteins offers significant potential for gaining mechanistic insights into biological systems and introduces new avenues for treating diseases, enhancing agriculture, and safeguarding the environment. While large protein language models (PLMs)…
Shadi Zabad, Yue Li, Simon Gravel
With the increasing availability of high quality genomic data from diverse cohorts, polygenic scores (PRS) have become a mainstay of genetic analyses of complex traits and diseases. Despite their proliferation in numerous research domains, a major obstacle to wider adoption in clinical settings has been the…
Yuxi Liu, Zhenhao Zhang, Mufan Qiu, Song Wang + 5 more
Single-cell RNA sequencing (scRNA-seq) enables high-resolution characterization of cellular heterogeneity, but its rich, complementary structure across cells and genes remains underexploited, especially in the presence of technical noise and sparsity. Effectively leveraging this multi-scale structure is essentially an…
Tingting Chen, Hongming Li, Hao Zheng, Yong Fan
Characterizing brain dynamic functional connectivity (dFC) patterns from functional Magnetic Resonance Imaging (fMRI) data is of paramount importance in imaging neuroscience and medicine. Recently, many graph neural network (GNN) models, combined with transformers or recurrent neural networks (RNNs), have shown great…
Yijingxiu Lu, Sangseon Lee, Soosung Kang, Sun Kim
In recent years, numerous deep learning models have been developed for drug-target interaction (DTI) prediction. These DTI models specialize in handling data with distinct distributions and features, often yielding inconsistent predictions when applied to unseen data points. This inconsistency poses a challenge for…
Liang Wang
How do multi-modal large language models that jointly process natural language and biological sequences (DNA, protein, structural alphabets) actually answer biological questions, especially sequence-grounded questions whose answer depends on residue-level patterns rather than literature recall? We introduce OmniGene-4…
Uthsav Chitra, Shu Dan, Fenna Krienen, Benjamin J. Raphael
Gene expression varies across a tissue due to both the organization of the tissue into spatial domains, i.e. discrete regions of a tissue with distinct cell type composition, and continuous spatial gradients of gene expression within different spatial domains. Spatially resolved transcriptomics (SRT) technologies…
Michael Huang, Yue Li
Advancements in single-cell transcriptomics methods have resulted in a wealth of single-cell RNA sequencing (scRNA-seq) data. Methods to learn cell representation from atlas-level scRNA-seq data across diverse tissues can shed light into cell functions implicated in diseases such as cancer. However, integrating…
Laetitia Meng-Papaxanthos, Ran Zhang, Gang Li, Marco Cuturi + 2 more
Modality matching in single-cell omics data analysis—i.e., matching cells across data sets collected using different types of genomic assays—has become an important problem, because unifying perspectives across different technologies holds the promise of yielding biological and clinical discoveries. However…
Nicholas Ho, Caleb N. Ellington, Jinyu Hou, Sohan Addagudi + 8 more
Developing a unified model of cellular systems is a canonical challenge in biology. Recently, a wealth of public single-cell RNA sequencing data as well as rapid scaling of self-supervised learning methods have provided new avenues to address this longstanding challenge. However, rapid parameter scaling has been…
Farhad Zamani, Asta Mannstaedt Rasmussen, Viktoria Schuster, Mathilde Hartvig Diekema + 2 more
MicroRNAs (miRNAs) are important post-transcriptional regulators, yet their expression is typically unobserved in single-cell and most bulk RNA-seq datasets. We present miDGD, a deep generative decoder model that predicts miRNA abundance directly from gene expression alone. Trained on bulk and single-cell datasets from…
Linxing Preston Jiang, Shirui Chen, Emmanuel Tanumihardja, Xiaochuang Han + 3 more
A key challenge in analyzing neuroscience datasets is the profound variability they exhibit across sessions, animals, and data modalities–i.e., heterogeneity. Several recent studies have demonstrated performance gains from pretraining neural foundation models on multi-session datasets, seemingly overcoming this…
Xinyu Jiang, Chenfei Ma, Kianoush Nazarpour
Myoelectric control systems translate electromyographic (EMG) signals into control commands, enabling immersive human-robot interactions in the real world and the Metaverse. The variability of EMG due to various confounding factors leads to significant performance degradation. Such variability can be mitigated by…