24 papers · ranked by Valyu relevance
Alexander Glandon, Khan M. Iftekharuddin
—Multisource domain adaptation (MDA) aims to use multiple source datasets with available labels to infer labels on a target dataset without available labels for target supervision. Prior works on MDA in the literature is ad-hoc as the pretraining of source models is either based on weight sharing or uses independently…
Sully F. Chen, Robert J. Steele, Glen M. Hocky, Beakal Lemeneh + 3 more
The transformer architecture has revolutionized bioinformatics and driven progress in the understanding and prediction of the properties of biomolecules. To date, most biosequence transformers have been trained on single-omic data-either proteins or nucleic acids-and have seen incredible success in downstream tasks in…
Authors not listed
Equivariant graph neural networks have shown remarkable success in molecular property prediction, but their performance on novel molecular geometries remains limited without extensive training data. We present a computationally efficient approach to cross-geometry pretraining for molecular systems that improves…
Qingyue Zhang, Changyong Chu, Haohao Fu, Tianren Peng + 4 more
—Transfer learning plays a vital role in improving model performance in data-scarce scenarios. However, naive uniform transfer from multiple source tasks may result in negative transfer, highlighting the need to properly balance the contributions of heterogeneous sources. Moreover, existing transfer learning methods…
Mary Isabelle Wisell, Nicholas Jacobs, Aayush Manandhar, Salimeh Yasaei Sekeh
Multi-source transfer learning faces a fundamental scalability bottleneck: existing approaches require either loading all K source models into memory simultaneously during parameter fusion, requiring O(K) memory, or deploying all models at inference time, making production deployment infeasible. We propose GRASP…
Shuo Yang, Bin Zhou, Yanjiang Wang, Weifeng Liu
To identify various oil well working conditions more accurately and practically from massive image data collected by multiple measured information sources of sucker-rod pumping wells, this paper proposes a working condition recognition method with three key aspects: curvelet pooling optimization technology…
Yajie Li, Qichang Zhao, Jianxin Wang
Accurate prediction of molecular properties is essential for accelerating drug discovery. While graph neural networks (GNNs) have achieved impressive progress, most models rely on atomic graphs with limited chemical semantics, leading to suboptimal generalization across diverse biochemical tasks. Pretraining alleviates…
Gabriel Bianchin de Oliveira, Fahad Saeed
Foundational models that learn the “language” of molecules are essential for accelerating material and drug discovery. These self-learning models can be trained on large collections of unlabelled molecules, enabling applications such as property prediction, molecule design, and screening for specific functions.…
Alex O. Davies, Riku Green, Telmo M. Silva Filho, Nirav Ajmeri
The principal benefit of unsupervised representation learning is that a pre-trained model can be fine-tuned where data or labels are scarce. Existing approaches for graph representation learning are domain specific, maintaining consistent node and edge features across the pre-training and target datasets. This has…
Bowen Li, Ziqiang Liu, Zhen Wang, Zhenyu Xu + 3 more
Advances in single-cell multimodal profiling have enabled a more systematic analysis of cellular biology, yet the rapid accumulation of large-scale, heterogeneous datasets poses substantial challenges for integrative analysis. Recently, Transformer-based cell language models (CLMs) are becoming powerful foundational…
Maximilien Burq, Peter Cimermancic, Charlie Kim, Dejan Stepec
Leading cellular foundation models have been trained on hundreds of millions of single-cell transcriptomes, with progress increasingly driven by larger datasets and model scaling. Here, we asked whether adding a proteomics modality can improve gene-level and cell-level representations beyond scaling RNA-only models. We…
Kuan Pang, Yanay Rosen, Kasia Kedzierska, Ziyuan He + 4 more
Biology emerges from interactions across physical scales, where molecular interactions drive cellular states, which in turn orchestrate multicellular tissue functions that collectively define health and disease. However, current computational models are often constrained to single scales in isolation, failing to…
Aidan Dempster, Brokoslaw Laschowski
A grand challenge in brain decoding is to develop algorithms that generalize across multiple subjects and tasks. Here, we developed a new computational framework to minimize negative transfer for domain-adaptive brain decoding by reframing source selection as a mixture model parameter estimation problem, allowing each…
Eduardo González-García, Albert J. Markvoort, Nadia A. Erkamp, Tom F. A. de Greef
Self-driving laboratories accelerate materials discovery by autonomously designing and executing experiments through closed-loop integration of robotics and artificial intelligence. Active learning with Gaussian processes has enabled efficient phase diagram mapping, reducing required measurements by approximately 80%…
Niccolò McConnell, Pardeep Vasudev, Daisuke Yamada, Daryl Cheng + 11 more
Background: Low-dose computed tomography (LDCT) employed in lung cancer screening (LCS) programmes is increasing in uptake worldwide. LCS programmes herald a generational opportunity to simultaneously detect cancer and non-cancer-related early-stage lung disease, yet these efforts are hampered by a shortage of…
Authors not listed
Accurate prediction of chemical reaction yields remains essential for accelerating synthesis optimization, yet current machine learning models face critical limitations in capturing temporal dynamics, providing calibrated uncertainty estimates, and explicitly modeling reactant-to-product transformations. Here we…
Mengyuan Zhao, Xinyue Tang, Jiawei Li, Cheng Liang + 3 more
Precise prediction of perturbation responses is essential in systems biology research, as it plays a pivotal role in characterizing cellular identities and elucidating the regulatory mechanisms of biological pathways. Existing perturbation-responses prediction approaches are predominantly confined to single-modality…
Qinglei Jiang, Tielin Shi, Xiuqun Hou, Biqi Miao + 7 more
Domain adaptation methods have been extensively studied for rolling bearing fault diagnosis under various conditions. However, some existing methods only consider the one-way embedding of original space into a low-dimensional subspace without backward validation, which leads to inaccurate embeddings of data and poor…
Haowen Wang, Yaxin Du, Jian Yang, Jiajun Wu + 7 more
Mid-training has become an important stage in modern LLM development, using large-scale curated mixtures to strengthen capabilities before final post-training. Its data selection problem is distinct: the data are optimized under a pretraining-style objective at near-pretraining scale, but are curated toward downstream…
Cormac Cureton, Narges Armanfard
Prior-Data Fitted networks (PFNs) have been very successful in tabular contexts, handling prediction tasks in context. However, they are designed for single-task inference, meaning that predicting several target values within a context requires repeated forward calls and precludes inter-task information sharing. We…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Authors not listed
Recent years have seen a growing interest in machine learning approaches for chemical tasks. The best existing methods focus on building base models that combine molecular graphs (“2D structures”) with atomic coordinates in 3D to predict molecular properties, typically through pre-training followed by fine-tuning on…
Authors not listed
Recent advances in machine learning force fields (MLFF) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference datasets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and…
Yan Gao, Yan Cui
Large-scale clinical and biomedical datasets increasingly contain both diverse subgroup attributes (e.g., demographic or clinical subgroups) and multiple prediction targets. Although various machine learning approaches can address subgroup differences or multi-target prediction, they often consider these aspects…