13 papers · ranked by Valyu relevance
Muhammad Hammad, Önder Babur, Hamid Abdul Basit, Mark van den Brand + 1 more
'Yilun Shang'] Software developers frequently reuse source code from repositories as it saves development time and effort. Code clones (similar code fragments) accumulated in these repositories represent often repeated functionalities and are candidates for reuse in an exploratory or rapid development. To facilitate…
Andrea Gurioli, Maurizio Gabbrielli, Stefano Zacchiroli, Stefan Wagner
'Stefan Wagner'] Code stylometry is the application of stylometry techniques to determine the authorship of software source code snippets. It is used in the industry to address use cases like plagiarism detection, code audits, and code review assignments. Most works in the code stylometry literature use machine…
Qihong Song, Haize Hu, Tebo Dai
Code search aims to search for code snippets from large codebase that are semantically related to natural query statements. Deep learning is a valuable method for solving code search tasks in which the quality of training data directly impacts the performance of deep-learning models. However, most existing…
Zhenhao Li, Hang Lei, Zhichao Ma, Fengyun Zhang + 4 more
'Yongpan Sheng' 'Hao Wang' 'Junyang Chen'] The code of industrial management software typically features few system API calls and a high number of customized variables and structures. This makes the similarity of such codes difficult to compute using text features or traditional neural network methods. In this paper…
Marc Joiret, Marine Leclercq, Gaspard Lambrechts, Francesca Rapino + 3 more
'Pierre Close' 'Gilles Louppe' 'Liesbet Geris'] The genetic code is textbook scientific knowledge that was soundly established without resorting to Artificial Intelligence (AI). The goal of our study was to check whether a neural network could re-discover, on its own, the mapping links between codons and amino acids…
Leonardo Costa Ribeiro, Américo Tristão Bernardes, Heliana Mello, Ramona Bongelli
'Ramona Bongelli'] Natural Language Processing (NLP) makes use of Artificial Intelligence algorithms to extract meaningful information from unstructured texts, i.e., content that lacks metadata and cannot easily be indexed or mapped onto standard database fields. It has several applications, from sentiment analysis and…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Kristína Machová, Marián Mach, Michal Porezaný, Seongsoo Cho
This article focuses on the problem of detecting disinformation about COVID-19 in online discussions. As the Internet expands, so does the amount of content on it. In addition to content based on facts, a large amount of content is being manipulated, which negatively affects the whole society. This effect is currently…
Jianhui Zeng, Zhiheng Qu, Bo Cai, Boris Ryabko
Source code summarization focuses on generating qualified natural language descriptions of a code snippet (e.g., functionality, usage and version). In an actual development environment, descriptions of the code are missing or not consistent with the code due to human factors, which makes it difficult for developers to…
Emrecan Kutay, Aylin Yener, Kai Niu, Meixia Tao + 1 more
This paper investigates point-to-point multimodal digital semantic communications in a task-oriented setup, where messages are classified at the receiver. We employ a pre-trained transformer model to extract semantic information and propose three methods for generating semantic codewords. First, we propose semantic…
Muhammad Zulqarnain, Ahmed Khalaf Zager Alsaedi, Rozaida Ghazali, Muhammad Ghulam Ghouse + 3 more
'Muhammad Ghulam Ghouse' 'Wareesa Sharif' 'Noor Aida Husaini' 'Abdel Hamid Soliman'] Question classification is one of the essential tasks for automatic question answering implementation in natural language processing (NLP). Recently, there have been several text-mining issues such as text classification, document…
Cristian Robledo, Francesca Sallicati, Gaël de Chalendar, Marcos Fernández + 4 more
'Marcos Fernández' 'Pablo de Castro' 'Eduardo Martín' 'Javier Gutiérrez' 'Yannis Bouachera'] This paper aims to introduce the innovative work carried out in the Horizon 2020 DECODER project - acronym for “DEveloper COmpanion for Documented and annotatEd code Reference” - (Grant Agreement no. 824231) by linking the…
Anandan Chinnalagu, Ashok Kumar Durairaj, Jude Duraisamy
Customer satisfaction and their positive sentiments are some of the various goals for successful companies. However, analyzing customer reviews to predict accurate sentiments have been proven to be challenging and time-consuming due to high volumes of collected data from various sources. Several researchers approach…