Search · four archives
Search · four archives
28 papers · ranked by Valyu relevance
Steve Nwaiwu
The successful application of large-scale transformer models in Natural Language Processing (NLP) is often hindered by the substantial computational cost and data requirements of full fine-tuning. This challenge is particularly acute in low-resource settings, where standard fine-tuning can lead to catastrophic…
Han Yuan, Johannes Linder, David R Kelley
DNA sequence deep learning models accurately predict epigenetic and transcriptional profiles, enabling analysis of gene regulation and genetic variant effects. While large-scale training models like Enformer and Borzoi are trained on abundant data, they cannot cover all cell states and assays, necessitating new model…
Olesya Razuvayevskaya, Ben Wu, João A. Leite, Freddy Heppell + 5 more
'Ivan Srba' 'Carolina Scarton' 'Kalina Bontcheva' 'Xingyi Song' 'Mario Graff-Guerrero'] Adapters and Low-Rank Adaptation (LoRA) are parameter-efficient fine-tuning techniques designed to make the training of language models more efficient. Previous results demonstrated that these methods can even improve performance on…
Feng Guan, Hao Hong, Yong Wang
In recent years, despite the remarkable performance of large-scale vision language models across various visual classification tasks, their substantial parameter counts and high fine-tuning costs have hindered deployment in resource-constrained cultural and artwork settings. This work specifically addresses the task of…
Nusrat Jahan Prottasha, Asif Mahmud, Md. Shohanur Islam Sobuj, Prakash Bhat + 3 more
'Prakash Bhat' 'Md Kowsher' 'Niloofar Yousefi' 'Ozlem Ozmen Garibay'] Large Language Models (LLMs) are gaining significant popularity in recent years for specialized tasks using prompts due to their low computational cost. Standard methods like prefix tuning utilize special, modifiable tokens that lack semantic meaning…
Marcos Treviso, Tianchu Ji, Ji-Ung Lee, Betty van Aken + 14 more
'Qingqing Cao' 'Manuel R. Ciosici' 'Michael Hassid' 'Kenneth Heafield' 'Sara Hooker' 'Pedro H. Martins' 'André F. T. Martins' 'Peter Milder' 'Colin Raffel' 'Edwin Simpson' 'Noam Slonim' 'Niranjan Balasubramanian' 'Leon Derczynski' 'Roy Schwartz'] Recent work in natural language processing (NLP) has yielded appealing…
Shiyun Xu, Zhiqi Bu
Parameter-efficient fine-tuning (PEFT) is a highly effective approach for adapting large pre-trained models to downstream tasks with minimal computational overhead. At the core, PEFT methods freeze most parameters and only trains a small subset (say < 0.1% of total parameters). Notably, different PEFT methods select…
Zhengxiang Shi, Aldo Lipani
Prompt tuning (PT), where a small amount of trainable soft (continuous) prompt vectors is affixed to the model input, has shown promising results across various tasks and model architecture for parameter-efficient fine-tuning (PEFT). PT stands out from other PEFT approaches because it maintains competitive performance…
Nikhil J. Dhinagar, Saket S. Ozarkar, Ketaki U. Buwa, Sophia I. Thomopoulos + 10 more
Recent innovations in artificial intelligence (AI) have increasingly focused on large-scale foundational models that are more general purpose in contrast to conventional models trained to perform specialized tasks. Transformer-based architectures have become the standard backbone in foundation models across data…
Samuel Sledzieski, Meghana Kshirsagar, Minkyung Baek, Bonnie Berger + 2 more
Proteomics has been revolutionized by large pre-trained protein language models, which learn unsupervised representations from large corpora of sequences. The parameters of these models are then fine-tuned in a supervised setting to tailor the model to a specific downstream task. However, as model size increases, the…
Zihao Fu, Haoran Yang, Anthony Man–Cho So, Wai Lam + 2 more
'Nigel Collier'] Fine-tuning pre-trained models has been ubiquitously proven to be effective in a wide range of NLP tasks. However, fine-tuning the whole model is parameter inefficient as it always yields an entirely new model for each task. Currently, many research works propose to only fine-tune a small portion of…
Kaibin Wei, Jianqiang Jing, Jiawei Liu, Qing Liu + 2 more
Time series forecasting models often face challenges in cross-domain fine-tuning, such as high training costs and limited adaptability. To address these issues, we propose a Cue-driven Feature Fusion Network (CFF-Net), which combines semantic cues from textual prompts with numerical time series features for…
Seda Bayat Toksöz, Gültekin Işık, Gökhan Şahin, Erdal Akin + 2 more
Automated visual inspection of photovoltaic (PV) cells is an important component of solar-module quality assurance. However, adapting modern pre-trained vision backbones to PV defect classification remains challenging because full fine-tuning requires substantial memory, naturally imbalanced datasets can reduce…
Shuai Zeng, Duolin Wang, Dong Xu
Signal peptides (SP) play a crucial role in protein translocation in cells. The development of large protein language models (PLMs) provides a new opportunity for SP prediction, especially for the categories with limited annotated data. We present a Parameter-Efficient Fine-Tuning (PEFT) framework for SP prediction…
Danilo Vucetic, Mohammadreza Tayaranian, Maryam Ziaeefard, James J. Clark + 2 more
'James J. Clark' 'Brett H. Meyer' 'Warren J. Gross'] Fine-tuning BERT-based models is resourceintensive in memory, computation, and time. While many prior works aim to improve inference efficiency via compression techniques, e.g., pruning, these works do not explicitly address the computational challenges of training…
Dong Yang, M. Amiri, Tejaswini Pedapati, Subhajit Chaudhury + 1 more
'Pin‐Yu Chen'] Fine-tuning large language models (LLMs) for downstream tasks has become increasingly crucial due to their widespread use and the growing availability of open-source models. However, the high memory costs associated with fine-tuning remain a significant challenge, especially as models increase in size.…
Authors not listed
Large Language Models (LLMs) based on transformer architectures excel at internet-scale tasks. However, real-world scientific scenarios—such as synthetic chemistry laboratories and autonomous experimental setups—typically involve incremental data generation in batches as new chemical reactions are conducted, unlike…
Reza Arabpour, Haitz Sáez de Ocáriz Borde, Anastasis Kratsios
Low-Rank Adapters (LoRAs) have transformed the fine-tuning of Large Language Models (LLMs) by enabling parameter-efficient updates. However, their widespread adoption remains limited by the reliance on GPU-based training. In this work, we propose a theoretically grounded approach to LoRA fine-tuning designed…
Nur Bengisu Çam, Hasin Rehana, Jie Zheng, Benu Bansal + 3 more
Protein-protein interactions (PPIs) play a crucial role in various biological processes, and understanding these interactions is essential for advancing biomedical research. Automated extraction and analysis of PPI information from the rapidly growing scientific literature remains an important challenge. We present a…
Muhammed Hasan Çelik, Xiaohui Xie
Title: Summary Protein language models (PLMs) have shown great promise in protein structure and function predictions, but their adoption is limited by computational cost. We address this challenge by enhancing the efficiency of evolutionary scale modeling (ESM). Using FlashAttention and sequence packing, we achieve…
Dominic Rufa, Joshua Fass, John D. Chodera
Alchemical free energy methods using molecular mechanics (MM) force fields are essential tools for predicting thermodynamic properties of small molecules, especially via free energy calculations that can estimate quantities relevant for drug discovery such as affinities, selectivities, the impact of target mutations…
Robert Schmirler, Michael Heinzinger, Burkhard Rost
Prediction methods inputting embeddings from protein Language Models (pLMs) have reached or even surpassed state-of-the-art (SOTA) performance on many protein prediction tasks. In natural language processing (NLP) fine-tuning large Language Models (LLMs) has become the de facto standard. In contrast, most pLM-based…
Authors not listed
The ability to generate crystal structures directly from textual descriptions marks a pivotal advancement in materials informatics and underscores the emerging role of large language models (LLMs) in inverse design. In this work, we introduce CrysText, a text-conditioned framework that generates crystal structures in…
Michael Oliver, G. Wang
This paper addresses the challenges of efficiently fine-tuning large language models (LLMs) by exploring data efficiency and hyperparameter optimization. We investigate the minimum data required for effective fine-tuning and propose a novel hyperparameter optimization method that leverages early-stage model…
David Buterez, Jon Paul Janet, Steven Kiddle, Pietro Liò
We investigate the potential of graph neural networks for transfer learning and improving molecular property prediction on sparse and expensive to acquire high-fidelity data by leveraging low-fidelity measurements as an inexpensive proxy for a targeted property ofinterest. This problem arises in discovery processes…
Authors not listed
Inverse molecular design aims to generate novel chemical structures that satisfy multiple property constraints, yet reinforcement-learning (RL) fine-tuning can be sensitive to how objectives are converted into a scalar reward. Here, we systematically analyze how scalarization choices and stabilization mechanisms shape…
Nathan Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi + 3 more
Massive scale, both in terms of data availability and computation, enables significant breakthroughs in key application areas of deep learning such as natural language processing (NLP) and computer vision. There is emerging evidence that scale may be a key ingredient in scientific deep learning, but the importance of…
Sterling Baird, Jason R. Hall, Taylor D. Sparks
Would you rather search for a line inside a cube or a point inside a square? This type of solution degeneracy often exists in physics-based simulations and wet-lab experiments, but constraining these degeneracies is often unsupported or difficult to implement in many optimization packages, requiring additional time and…