14 papers · ranked by Valyu relevance
Vasco Faria, António Rito Silva
—Migrating a monolith application into a microservices architecture can benefit from automation methods, which speed up the migration and improve the decomposition results. One of the current approaches that guide software architects on the migration is to group monolith domain entities into microservices, using the…
Yutao Xie, Jiayi Lin, Hande Dong, Lei Zhang + 1 more
Code writing is repetitive and predictable, inspiring us to develop various code intelligence techniques. This survey focuses on code search, that is, to retrieve code that matches a given natural language query by effectively capturing the semantic similarity between the query and code. Deep learning, being able to…
Ruibo Shi, Lili Tao, Rohan Saphal, Fran Silavong + 1 more
We present CV4Code, a compact and effective computer vision method for sourcecode understanding. Our method leverages the contextual and the structural information available from the code snippet by treating each snippet as a two-dimensional image, which naturally encodes the context and retains the underlying…
Peter Samoaa, Mehrdad Vasheghani Farahani, Antonio Longa, Philipp Leitner + 1 more
Tasks Authors: ['Peter Samoaa' 'Mehrdad Vasheghani Farahani' 'Antonio Longa' 'Philipp Leitner' 'Morteza Haghir Chehreghani'] Abstract—The landscape of deep learning has vastly expanded the frontiers of source code analysis, particularly through the utilization of structural representations such as Abstract Syntax Trees…
Sairamvinay Vijayaraghavan, Jinxiao Song, David A. Tomassi, Siddhartha Punj + 1 more
Classification of text is an important field of research and a core task in Natural Language Processing (NLP). It spans many different domains from determining "fake" news, finding spam emails, and language detection. A common problem in software development is generating the appropriate code snippet for a task. There…
Dharma KC, Clayton T. Morrison
Neural machine translation (NMT) methods developed for natural language processing have been shown to be highly successful in automating translation from one natural language to another. Recently, these NMT methods have been adapted to the generation of program code. In NMT for code generation, the task is to generate…
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Md Rafiqul Islam Rabin, Mohammad Amin Alipour
—There are several approaches for encoding source code in the input vectors of neural models. These approaches attempt to include various syntactic and semantic features of input programs in their encoding. In this paper, we investigate CODE2SNAPSHOT, a novel representation of the source code that is based on the…
Zhiwei Xu, Min Zhou, Xibin Zhao, Yang Chen + 2 more
Code representations (a.k.a., embeddings) is of great importance in deep learning-based software engineering techniques. A highquality representation model can significantly improve the performance of many downstream tasks, such as code search [13, 23, 42], code clone detection [21, 48, 54], and bug localization [27].…
Erfan Al-Hossami, Samira Shaikh
In this survey paper, we overview major deep learning methods used in Natural Language Processing (NLP) and source code over the last 35 years. Next, we present a survey of the applications of Artificial Intelligence (AI) for source code, also known as Code Intelligence (CI) and Programming Language Processing (PLP).…
Anthony Varkey, Siyuan Jiang, Weijing Huang
Embeddings Authors: ['Anthony Varkey' 'Siyuan Jiang' 'Weijing Huang'] Abstract—Pretrained language models for code token embeddings are used in code search, code clone detection, and other code-related tasks. Similarly, code function embeddings are useful in such tasks. However, there is no out-of-box models for…
Changan Niu, Chuanyi Li, Vincent Ng, Jidong Ge + 2 more
Pre-Training has revolutionized the way computational models are trained in the natural language processing (NLP) community [12, 13, 43, 44]. For a long time, supervised learning has been the most successful natural language learning paradigm. The pioneers of the pre-training idea challenged this view by showing that a…
Ensheng Shi, Wenchao Gub, Yanlin Wang, Lun Du + 4 more
'Han Shi' 'Dongmei Zhang' 'Hongbin Sun'] Abstract—Code search aims to retrieve semantically relevant code snippets for a given natural language query. Recently, many approaches employing contrastive learning have shown promising results on code representation learning and greatly improved the performance of code…
Frank F. Xu, Uri Alon, Graham Neubig, Vincent J. Hellendoorn
Large language models (LMs) of code have recently shown tremendous promise in completing code and synthesizing code from natural language descriptions. However, the current state-of-the-art code LMs (e.g., Codex (Chen et al., 2021)) are not publicly available, leaving many questions about their model and data design…