14 papers · ranked by Valyu relevance
Li HaoChen, Chunyan Miao, Cyril Leung, Yanxian Huang + 3 more
'Hongyu Zhang' 'Yanlin Wang'] Code search, which aims at retrieving the most relevant code fragment for a given natural language query, is a common activity in software development practice. Recently, contrastive learning is widely used in code search research, where many data augmentation approaches for source code…
Zeming Dong, Qiang Hu, Yuejun Guo, Maxime Cordy + 3 more
'Yves Le Traon' 'Jianjun Zhao'] Abstract—Inspired by the great success of Deep Neural Networks (DNNs) in natural language processing (NLP), DNNs have been increasingly applied in source code analysis and attracted significant attention from the software engineering community. Due to its data-driven nature, a DNN model…
Zeming Dong, Qiang Hu, Xiaofei Xie, Maxime Cordy + 2 more
'Jianjun Zhao'] Pre-trained code models lead the era of code intelligence. Many models have been designed with impressive performance recently. However, one important problem, data augmentation for code data that automatically helps developers prepare training data lacks study in the field of code learning. In this…
Joonghyuk Hahn, Hyeseon Ahn, Jungin Kim, Soohan Lim + 1 more
Time complexity is a theoretic measure to determine the amount of time the algorithm needs for its execution. In reality, developers write algorithms into code snippets within limited resources, making the calculation of a code's time complexity a fundamental task. However, determining the precise time complexity of a…
Yanlin Wang, Lianghong Guo, Ensheng Shi, Wenqing Chen + 7 more
'Wanjun Zhong' 'Menghan Wang' 'Hui Li' 'Hongyu Zhang' 'Ziyu Lyu' 'Zibin Zheng'] Abstract—Code search plays a crucial role in software development, enabling developers to retrieve and reuse code using natural language queries. While the performance of code search models improves with an increase in high-quality data…
Dong Li, Yelong Shen, Ruoming Jin, Yi Mao + 2 more
Pre-trained language models have achieved promising success in code retrieval tasks, where a natural language documentation query is given to find the most relevant existing code snippet. However, existing models focus only on optimizing the documentation code pairs by embedding them into latent space, without the…
Pinzhen Chen, Γεράσιμος Λάμπουρας
Advances in natural language processing, such as transfer learning from pre-trained language models, have impacted how models are trained for programming language tasks too. Previous research primarily explored code pre-training and expanded it through multi-modality and multi-tasking, yet the data for downstream tasks…
Zeming Dong, Qiang Hu, Yuejun Guo, Zhenya Zhang + 4 more
'Mike Papadakis' 'Yves Le Traon' 'Jianjun Zhao'] The next era of program understanding is being propelled by the use of machine learning to solve software problems. Recent studies have shown surprising results of source code learning, which applies deep neural networks (DNNs) to various critical software tasks, e.g.…
Cristina Improta, Pietro Liguori, Roberto Natella, Bojan Čukić + 1 more
'Domenico Cotroneo'] In this work, we present a method to add perturbations to the code descriptions to create new inputs in natural language (NL) from well-intentioned developers that diverge from the original ones due to the use of new words or because they miss part of them. The goal is to analyze how and to what…
Shreya Shukla, Prajwal Gatti, Yogesh Kumar, Vikash Yadav + 1 more
'Anand Mishra'] Abstract. Computer programming textbooks and software documentations often contain flowcharts to illustrate the flow of an algorithm or procedure. Modern OCR engines often tag these flowcharts as graphics and ignore them in further processing. In this paper, we work towards making flowchart images…
Zezhou Yang, Sirong Chen, Cuiyun Gao, Zhenhao Li + 3 more
and Opportunities Authors: ['Zezhou Yang' 'Sirong Chen' 'Cuiyun Gao' 'Zhenhao Li' 'Xing Hu' 'Kui Liu' 'Xin Xia'] Code generation aims to automatically generate code snippets of specific programming language according to natural language descriptions. The continuous advancements in deep learning, particularly…
Alex Sheng
Recent progress in large-scale language models has enabled breakthroughs in previously intractable computer programming tasks. Prior work in meta-learning and neural architecture search has led to substantial successes across various task domains, spawning myriad approaches for algorithmically optimizing the design and…
Wen Qi, Jiahao Cao, Debasis Poddar, Sophia Li + 1 more
Semantic-Preserving Data Augmentation Authors: ['Wen Qi' 'Jiahao Cao' 'Debasis Poddar' 'Sophia Li' 'Xinda Wang'] Abstract. With the rapid development and widespread use of advanced network systems, software vulnerabilities pose a significant threat to secure communications and networking. Learning-based vulnerability…
Sahil Suneja, Yufan Zhuang, Yunhui Zheng, Jim Laredo + 1 more
'Alessandro Morari'] AI modeling for source code understanding tasks has been making significant progress, and is being adopted in production development pipelines. However, reliability concerns, especially whether the models are actually learning task-related aspects of source code, are being raised. While recent…