14 papers · ranked by Valyu relevance
Authors not listed
Vulnerability code-bases often suffer from severe imbalance, limiting the effectiveness of Deep Learning-based vulnerability classifiers. Data Augmentation could help solve this by mitigating the scarcity of under-represented CWEs. In this context, we investigate LLM-based augmentation for vulnerable functions…
Li HaoChen, Chunyan Miao, Cyril Leung, Yanxian Huang + 3 more
'Hongyu Zhang' 'Yanlin Wang'] Code search, which aims at retrieving the most relevant code fragment for a given natural language query, is a common activity in software development practice. Recently, contrastive learning is widely used in code search research, where many data augmentation approaches for source code…
Zeming Dong, Qiang Hu, Xiaofei Xie, Maxime Cordy + 2 more
'Jianjun Zhao'] Pre-trained code models lead the era of code intelligence. Many models have been designed with impressive performance recently. However, one important problem, data augmentation for code data that automatically helps developers prepare training data lacks study in the field of code learning. In this…
Zeming Dong, Qiang Hu, Yuejun Guo, Maxime Cordy + 3 more
'Yves Le Traon' 'Jianjun Zhao'] Abstract—Inspired by the great success of Deep Neural Networks (DNNs) in natural language processing (NLP), DNNs have been increasingly applied in source code analysis and attracted significant attention from the software engineering community. Due to its data-driven nature, a DNN model…
Joonghyuk Hahn, Hyeseon Ahn, Jungin Kim, Soohan Lim + 1 more
Time complexity is a theoretic measure to determine the amount of time the algorithm needs for its execution. In reality, developers write algorithms into code snippets within limited resources, making the calculation of a code's time complexity a fundamental task. However, determining the precise time complexity of a…
Yanlin Wang, Lianghong Guo, Ensheng Shi, Wenqing Chen + 7 more
'Wanjun Zhong' 'Menghan Wang' 'Hui Li' 'Hongyu Zhang' 'Ziyu Lyu' 'Zibin Zheng'] Abstract—Code search plays a crucial role in software development, enabling developers to retrieve and reuse code using natural language queries. While the performance of code search models improves with an increase in high-quality data…
Mehdi Bahrami, N. C. Shrikanth, Yuji Mizobuchi, Lei Liu + 3 more
'Masahiro Fukuyori' 'Weipeng Chen' 'Kazuki Munakata'] Code retrieval is allowing software engineers to search codes through a natural language query, which relies on both natural language processing and software engineering techniques. There have been several attempts on code retrieval from searching snippet codes to…
Zeming Dong, Qiang Hu, Yuejun Guo, Zhenya Zhang + 4 more
'Mike Papadakis' 'Yves Le Traon' 'Jianjun Zhao'] The next era of program understanding is being propelled by the use of machine learning to solve software problems. Recent studies have shown surprising results of source code learning, which applies deep neural networks (DNNs) to various critical software tasks, e.g.…
Cristina Improta, Pietro Liguori, Roberto Natella, Bojan Čukić + 1 more
'Domenico Cotroneo'] In this work, we present a method to add perturbations to the code descriptions to create new inputs in natural language (NL) from well-intentioned developers that diverge from the original ones due to the use of new words or because they miss part of them. The goal is to analyze how and to what…
Shreya Shukla, Prajwal Gatti, Yogesh Kumar, Vikash Yadav + 1 more
'Anand Mishra'] Abstract. Computer programming textbooks and software documentations often contain flowcharts to illustrate the flow of an algorithm or procedure. Modern OCR engines often tag these flowcharts as graphics and ignore them in further processing. In this paper, we work towards making flowchart images…
Sahil Suneja, Yufan Zhuang, Yunhui Zheng, Jim Laredo + 1 more
'Alessandro Morari'] AI modeling for source code understanding tasks has been making significant progress, and is being adopted in production development pipelines. However, reliability concerns, especially whether the models are actually learning task-related aspects of source code, are being raised. While recent…
Zezhou Yang, Sirong Chen, Cuiyun Gao, Zhenhao Li + 3 more
and Opportunities Authors: ['Zezhou Yang' 'Sirong Chen' 'Cuiyun Gao' 'Zhenhao Li' 'Xing Hu' 'Kui Liu' 'Xin Xia'] Code generation aims to automatically generate code snippets of specific programming language according to natural language descriptions. The continuous advancements in deep learning, particularly…
Wen Qi, Jiahao Cao, Debasis Poddar, Sophia Li + 1 more
Semantic-Preserving Data Augmentation Authors: ['Wen Qi' 'Jiahao Cao' 'Debasis Poddar' 'Sophia Li' 'Xinda Wang'] Abstract. With the rapid development and widespread use of advanced network systems, software vulnerabilities pose a significant threat to secure communications and networking. Learning-based vulnerability…
Alex Sheng
Recent progress in large-scale language models has enabled breakthroughs in previously intractable computer programming tasks. Prior work in meta-learning and neural architecture search has led to substantial successes across various task domains, spawning myriad approaches for algorithmically optimizing the design and…