12 papers · ranked by Valyu relevance
Yutao Xie, Jiayi Lin, Hande Dong, Lei Zhang + 1 more
Code writing is repetitive and predictable, inspiring us to develop various code intelligence techniques. This survey focuses on code search, that is, to retrieve code that matches a given natural language query by effectively capturing the semantic similarity between the query and code. Deep learning, being able to…
Peter Samoaa, Mehrdad Vasheghani Farahani, Antonio Longa, Philipp Leitner + 1 more
Tasks Authors: ['Peter Samoaa' 'Mehrdad Vasheghani Farahani' 'Antonio Longa' 'Philipp Leitner' 'Morteza Haghir Chehreghani'] Abstract—The landscape of deep learning has vastly expanded the frontiers of source code analysis, particularly through the utilization of structural representations such as Abstract Syntax Trees…
Sairamvinay Vijayaraghavan, Jinxiao Song, David A. Tomassi, Siddhartha Punj + 1 more
Classification of text is an important field of research and a core task in Natural Language Processing (NLP). It spans many different domains from determining "fake" news, finding spam emails, and language detection. A common problem in software development is generating the appropriate code snippet for a task. There…
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Muhammad Hammad, Önder Babur, Hamid Abdul Basit, Mark van den Brand + 1 more
'Yilun Shang'] Software developers frequently reuse source code from repositories as it saves development time and effort. Code clones (similar code fragments) accumulated in these repositories represent often repeated functionalities and are candidates for reuse in an exploratory or rapid development. To facilitate…
Md Rafiqul Islam Rabin, Mohammad Amin Alipour
—There are several approaches for encoding source code in the input vectors of neural models. These approaches attempt to include various syntactic and semantic features of input programs in their encoding. In this paper, we investigate CODE2SNAPSHOT, a novel representation of the source code that is based on the…
Zhiwei Xu, Min Zhou, Xibin Zhao, Yang Chen + 2 more
Code representations (a.k.a., embeddings) is of great importance in deep learning-based software engineering techniques. A highquality representation model can significantly improve the performance of many downstream tasks, such as code search [13, 23, 42], code clone detection [21, 48, 54], and bug localization [27].…
Lili Liu, Zhen Li, Yu Wen, Penglong Chen + 1 more
Software vulnerabilities have led to system attacks and data leakage incidents, and software vulnerabilities have gradually attracted attention. Vulnerability detection had become an important research direction. In recent years, Deep Learning (DL)-based methods had been applied to vulnerability detection. The DL-based…
Andrea Gurioli, Maurizio Gabbrielli, Stefano Zacchiroli, Stefan Wagner
'Stefan Wagner'] Code stylometry is the application of stylometry techniques to determine the authorship of software source code snippets. It is used in the industry to address use cases like plagiarism detection, code audits, and code review assignments. Most works in the code stylometry literature use machine…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Kristína Machová, Marián Mach, Michal Porezaný, Seongsoo Cho
This article focuses on the problem of detecting disinformation about COVID-19 in online discussions. As the Internet expands, so does the amount of content on it. In addition to content based on facts, a large amount of content is being manipulated, which negatively affects the whole society. This effect is currently…
Jianhui Zeng, Zhiheng Qu, Bo Cai, Boris Ryabko
Source code summarization focuses on generating qualified natural language descriptions of a code snippet (e.g., functionality, usage and version). In an actual development environment, descriptions of the code are missing or not consistent with the code due to human factors, which makes it difficult for developers to…