13 papers · ranked by Valyu relevance
Rhys Compton, Eibe Frank, Panos Patros, Abigail Koay
Automatic source code analysis in key areas of software engineering, such as code security, can benefit from Machine Learning (ML). However, many standard ML approaches require a numeric representation of data and cannot be applied directly to source code. Thus, to enable ML, we need to embed source code into numeric…
Peter Samoaa, Mehrdad Vasheghani Farahani, Antonio Longa, Philipp Leitner + 1 more
Tasks Authors: ['Peter Samoaa' 'Mehrdad Vasheghani Farahani' 'Antonio Longa' 'Philipp Leitner' 'Morteza Haghir Chehreghani'] Abstract—The landscape of deep learning has vastly expanded the frontiers of source code analysis, particularly through the utilization of structural representations such as Abstract Syntax Trees…
Wei Ma, Mengjie Zhao, Ezekiel Soremekun, Qiang Hu + 5 more
'Mike Papadakis' 'Maxime Cordy' 'Xiaofei Xie' 'Yves Le Traon'] Code embedding is a keystone in the application of machine learning on several Software Engineering (SE) tasks. To effectively support a plethora of SE tasks, the embedding needs to capture program syntax and semantics in a way that is generic. To this end…
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Nghi D. Q. Bui, Yijun Yu, Lingxiao Jiang
—Learning code representations has found many uses in software engineering, such as code classification, code search, code comment generation, and bug prediction. Although representations of code in tokens, syntax trees, dependency graphs, paths in trees, or the combinations of their variants have been proposed…
Abdullah Al Ishtiaq, Masum Hasan, Md. Mahim Anjum Haque, Kazi Sajeed Mehrab + 4 more
'Kazi Sajeed Mehrab' 'Tanveer Muttaqueen' 'Tahmid Hasan' 'Anindya Iqbal' 'Rifat Shahriyar'] Millions of repetitive code snippets are submitted to code repositories every day. To search from these large codebases using simple natural language queries would allow programmers to ideate, prototype, and develop easier and…
Nghi D. Q. Bui, Yijun Yu, Lingxiao Jiang
We propose Corder, a self-supervised contrastive learning framework for source code model. Corder is designed to alleviate the need of labeled data for code retrieval and code summarization tasks. The pre-trained model of Corder can be used in two ways: (1) it can produce vector representation of code which can be…
Binita Rajbanshi, Anuj Guruacharya
Emerging generative models for biology focus on DNA, non-coding RNA, or proteins, ignoring information hidden in mRNA. Additionally, in protein engineering and mRNA therapeutics the design of mRNA sequences is still a challenge, lacking a clear framework. Here, we introduce and rigorously evaluate two novel methods: a…
Logan Hallee, Nikolaos Rafailidis, Jason P. Gleghorn
Recent advancements in Protein Language Models (pLMs) have enabled high-throughput analysis of proteins through primary sequence alone. At the same time, newfound evidence illustrates that codon usage bias is remarkably predictive and can even change the final structure of a protein. Here, we explore these findings by…
Zhangyang Gao, Cheng Tan, Stan Z. Li
Protein structure tokenization has attracted increasing attention in both protein representation learning and generation. While recent work, like FoldToken2 and ESM3, has achieved good reconstruction performance, the compressoin ratio is still limited. In this work, we propose FoldToken3, a novel protein structure…
Yanshuo Chen, Yuming Zhang, Joshua Li, Boxue Tian + 1 more
Codon optimization involves selecting synonymous codons to match host-specific preferences. It is critical for heterologous expression but remains challenging due to the combinatorial design space. Under long-term evolutionary selection, natural coding sequences are near-optimal compromises between translational…
Melissa Sanabria, Jonas Hirsch, Anna R. Poetsch
Large Language Models (LLMs) on natural language have achieved a level of performance that allows the generation of coherent and syntactically correct text. DNA sequence of genomes follows rules similar to natural language, but a distinguishing factor is the absence of a concept analogous to words. We established…
Jiayi Li, Hong-Sheng Lai, Litian Liang, Shiyi Du + 2 more
Codon sequence design is crucial for generating mRNA sequences with desired functional properties for tasks such as developing mRNA vaccines or gene editing therapies. Yet existing methods lack flexibility and controllability to adapt to various design objectives. We propose a novel machine learning-based framework…