13 papers · ranked by Valyu relevance
Daria Kryvosheieva, Saba Sturua, Michael Günther, Scott Martens + 1 more
jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically similar code snippets across programming languages. It makes innovative use of an autoregressive backbone pre-trained on both text and code…
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Zhao, Yu, Gong, Lina + 8 more
Vulnerability detection is garnering increasing attention in software engineering, since code vulnerabilities possibly pose significant security. Recently, reusing various code pre-trained models (e.g., CodeBERT, CodeT5, and CodeGen) has become common for code embedding without providing reasonable justifications in…
Benedikt Fein, Gordon Fraser
The trend of embedding source code for machine learning applications also enables new opportunities in learning analytics in programming education, but which code embedding approach is most suitable for learning analytics remains an open question. A common approach to embedding source code lies in treating the code as…
Anthony Varkey, Siyuan Jiang, Weijing Huang
Embeddings Authors: ['Anthony Varkey' 'Siyuan Jiang' 'Weijing Huang'] Abstract—Pretrained language models for code token embeddings are used in code search, code clone detection, and other code-related tasks. Similarly, code function embeddings are useful in such tasks. However, there is no out-of-box models for…
Ruibo Shi, Lili Tao, Rohan Saphal, Fran Silavong + 1 more
We present CV4Code, a compact and effective computer vision method for sourcecode understanding. Our method leverages the contextual and the structural information available from the code snippet by treating each snippet as a two-dimensional image, which naturally encodes the context and retains the underlying…
Saiteja Utpala, Alex Gu, Pin Yu Chen
Recently, code language models have achieved notable advancements in addressing a diverse array of essential code comprehension and generation tasks. Yet, the field lacks a comprehensive deep dive and understanding of the code embeddings of multilingual code models. In this paper, we present a comprehensive study on…
Ankit Kulshrestha, Vishwas Lele
There has been a steadily growing interest in development of novel methods to learn a representation of a given input data and subsequently using them for several downstream tasks. The field of natural language processing has seen a significant improvement in different tasks by incorporating pretrained embeddings into…
Hasnain Heickal, Andrew Lan
Providing effective feedback for programming assignments in computer science education can be challenging: students solve problems by iteratively submitting code, executing it, and using limited feedback from the compiler or the auto-grader to debug. Analyzing student debugging behavior in this process may reveal…
Yu Zhao, Lina Gong, Haoxiang Zhang, Yaoshen Yu + 1 more
Pre-trained language models have demonstrated powerful capabilities in the field of natural language processing (NLP). Recently, code pretrained model (PTM), which draw from the experiences of the NLP field, have also achieved state-of-the-art results in many software engineering (SE) downstream tasks. These code PTMs…
Zhiwei Xu, Min Zhou, Xibin Zhao, Yang Chen + 2 more
Code representations (a.k.a., embeddings) is of great importance in deep learning-based software engineering techniques. A highquality representation model can significantly improve the performance of many downstream tasks, such as code search [13, 23, 42], code clone detection [21, 48, 54], and bug localization [27].…
Ensheng Shi, Wenchao Gub, Yanlin Wang, Lun Du + 4 more
'Han Shi' 'Dongmei Zhang' 'Hongbin Sun'] Abstract—Code search aims to retrieve semantically relevant code snippets for a given natural language query. Recently, many approaches employing contrastive learning have shown promising results on code representation learning and greatly improved the performance of code…
Daria Cherniuk, Nikita Sukhorukov, Gusak, Danil + 5 more
Retrieval-augmented generation has emerged as one of the most effective approaches for code completion, particularly when context from a surrounding repository is essential. However, incorporating context significantly extends sequence length, leading to slower inference—a critical limitation for interactive settings…