14 papers · ranked by Valyu relevance
Daria Kryvosheieva, Saba Sturua, Michael Günther, Scott Martens + 1 more
jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically similar code snippets across programming languages. It makes innovative use of an autoregressive backbone pre-trained on both text and code…
José Cambronero, Hongyu Li, Seohyun Kim, Koushik Sen + 1 more
'Satish Chandra'] There have been multiple recent proposals on using deep neural networks for code search using natural language. Common across these proposals is the idea of embedding code and natural language queries into real vectors and then using vector distance to approximate semantic correlation between code and…
Zimin Chen, Martin Monperrus
Natural language processing has improved tremendously after the success of word embedding techniques such as word2vec. Recently, the same idea has been applied on source code with encouraging results. In this survey, we aim to collect and discuss the usage of word embedding techniques on programs and source code. The…
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Rhys Compton, Eibe Frank, Panos Patros, Abigail Koay
Automatic source code analysis in key areas of software engineering, such as code security, can benefit from Machine Learning (ML). However, many standard ML approaches require a numeric representation of data and cannot be applied directly to source code. Thus, to enable ML, we need to embed source code into numeric…
Wei Ma, Mengjie Zhao, Ezekiel Soremekun, Qiang Hu + 5 more
'Mike Papadakis' 'Maxime Cordy' 'Xiaofei Xie' 'Yves Le Traon'] Code embedding is a keystone in the application of machine learning on several Software Engineering (SE) tasks. To effectively support a plethora of SE tasks, the embedding needs to capture program syntax and semantics in a way that is generic. To this end…
Rafael Michael Karampatsis, Charles Sutton
Continuous embeddings of tokens in computer programs have been used to support a variety of software development tools, including readability, code search, and program repair. Contextual embeddings are common in natural language processing but have not been previously applied in software engineering. We introduce a new…
Jin Qin, Zihan Liao, Ziyin Zhang, Hang Yu + 2 more
We present C2LLM - Contrastive Code Large Language Models, a family of code embedding models in both 0.5B and 7B sizes. Building upon Qwen-2.5-Coder backbones, C2LLM adopts a Pooling by Multihead Attention (PMA) module for generating sequence embedding from token embeddings, effectively 1) utilizing the LLM's causal…
Abdullah Al Ishtiaq, Masum Hasan, Md. Mahim Anjum Haque, Kazi Sajeed Mehrab + 4 more
'Kazi Sajeed Mehrab' 'Tanveer Muttaqueen' 'Tahmid Hasan' 'Anindya Iqbal' 'Rifat Shahriyar'] Millions of repetitive code snippets are submitted to code repositories every day. To search from these large codebases using simple natural language queries would allow programmers to ideate, prototype, and develop easier and…
Wenchao Gu, Zongjie Li, Cuiyun Gao, Chaozheng Wang + 3 more
'Zenglin Xu' 'Michael R. Lyu'] Code retrieval is a common practice for programmers to reuse existing code snippets in the opensource repositories. Given a user query (i.e., a natural language description), code retrieval aims at searching the most relevant ones from a set of code snippets. The main challenge of…
Ruibo Shi, Lili Tao, Rohan Saphal, Fran Silavong + 1 more
We present CV4Code, a compact and effective computer vision method for sourcecode understanding. Our method leverages the contextual and the structural information available from the code snippet by treating each snippet as a two-dimensional image, which naturally encodes the context and retains the underlying…
Nghi D. Q. Bui, Yijun Yu, Lingxiao Jiang
We propose Corder, a self-supervised contrastive learning framework for source code model. Corder is designed to alleviate the need of labeled data for code retrieval and code summarization tasks. The pre-trained model of Corder can be used in two ways: (1) it can produce vector representation of code which can be…
Ensheng Shi, Wenchao Gub, Yanlin Wang, Lun Du + 4 more
'Han Shi' 'Dongmei Zhang' 'Hongbin Sun'] Abstract—Code search aims to retrieve semantically relevant code snippets for a given natural language query. Recently, many approaches employing contrastive learning have shown promising results on code representation learning and greatly improved the performance of code…
Hao Wang, Jia Zhang, Yingce Xia, Jiang Bian + 2 more
'Tie‐Yan Liu'] Abstract—Semantic code search, which aims to retrieve code snippets relevant to a given natural language query, has attracted many research efforts with the purpose of accelerating software development. The huge amount of online publicly available code repositories has prompted the employment of deep…