14 papers · ranked by Valyu relevance
Rhys Compton, Eibe Frank, Panos Patros, Abigail Koay
Automatic source code analysis in key areas of software engineering, such as code security, can benefit from Machine Learning (ML). However, many standard ML approaches require a numeric representation of data and cannot be applied directly to source code. Thus, to enable ML, we need to embed source code into numeric…
Peter Samoaa, Mehrdad Vasheghani Farahani, Antonio Longa, Philipp Leitner + 1 more
Tasks Authors: ['Peter Samoaa' 'Mehrdad Vasheghani Farahani' 'Antonio Longa' 'Philipp Leitner' 'Morteza Haghir Chehreghani'] Abstract—The landscape of deep learning has vastly expanded the frontiers of source code analysis, particularly through the utilization of structural representations such as Abstract Syntax Trees…
Wei Ma, Mengjie Zhao, Ezekiel Soremekun, Qiang Hu + 5 more
'Mike Papadakis' 'Maxime Cordy' 'Xiaofei Xie' 'Yves Le Traon'] Code embedding is a keystone in the application of machine learning on several Software Engineering (SE) tasks. To effectively support a plethora of SE tasks, the embedding needs to capture program syntax and semantics in a way that is generic. To this end…
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Nghi D. Q. Bui, Yijun Yu, Lingxiao Jiang
—Learning code representations has found many uses in software engineering, such as code classification, code search, code comment generation, and bug prediction. Although representations of code in tokens, syntax trees, dependency graphs, paths in trees, or the combinations of their variants have been proposed…
Abdullah Al Ishtiaq, Masum Hasan, Md. Mahim Anjum Haque, Kazi Sajeed Mehrab + 4 more
'Kazi Sajeed Mehrab' 'Tanveer Muttaqueen' 'Tahmid Hasan' 'Anindya Iqbal' 'Rifat Shahriyar'] Millions of repetitive code snippets are submitted to code repositories every day. To search from these large codebases using simple natural language queries would allow programmers to ideate, prototype, and develop easier and…
Nghi D. Q. Bui, Yijun Yu, Lingxiao Jiang
We propose Corder, a self-supervised contrastive learning framework for source code model. Corder is designed to alleviate the need of labeled data for code retrieval and code summarization tasks. The pre-trained model of Corder can be used in two ways: (1) it can produce vector representation of code which can be…
Muhammad Hammad, Önder Babur, Hamid Abdul Basit, Mark van den Brand + 1 more
'Yilun Shang'] Software developers frequently reuse source code from repositories as it saves development time and effort. Code clones (similar code fragments) accumulated in these repositories represent often repeated functionalities and are candidates for reuse in an exploratory or rapid development. To facilitate…
Andrea Gurioli, Maurizio Gabbrielli, Stefano Zacchiroli, Stefan Wagner
'Stefan Wagner'] Code stylometry is the application of stylometry techniques to determine the authorship of software source code snippets. It is used in the industry to address use cases like plagiarism detection, code audits, and code review assignments. Most works in the code stylometry literature use machine…
Qihong Song, Haize Hu, Tebo Dai
Code search aims to search for code snippets from large codebase that are semantically related to natural query statements. Deep learning is a valuable method for solving code search tasks in which the quality of training data directly impacts the performance of deep-learning models. However, most existing…
Marc Joiret, Marine Leclercq, Gaspard Lambrechts, Francesca Rapino + 3 more
'Pierre Close' 'Gilles Louppe' 'Liesbet Geris'] The genetic code is textbook scientific knowledge that was soundly established without resorting to Artificial Intelligence (AI). The goal of our study was to check whether a neural network could re-discover, on its own, the mapping links between codons and amino acids…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Kristína Machová, Marián Mach, Michal Porezaný, Seongsoo Cho
This article focuses on the problem of detecting disinformation about COVID-19 in online discussions. As the Internet expands, so does the amount of content on it. In addition to content based on facts, a large amount of content is being manipulated, which negatively affects the whole society. This effect is currently…
Jianhui Zeng, Zhiheng Qu, Bo Cai, Boris Ryabko
Source code summarization focuses on generating qualified natural language descriptions of a code snippet (e.g., functionality, usage and version). In an actual development environment, descriptions of the code are missing or not consistent with the code due to human factors, which makes it difficult for developers to…