13 papers · ranked by Valyu relevance
Qihong Song, Haize Hu, Tebo Dai
Code search aims to search for code snippets from large codebase that are semantically related to natural query statements. Deep learning is a valuable method for solving code search tasks in which the quality of training data directly impacts the performance of deep-learning models. However, most existing…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Fang-Yi Su, Gia-Han Ngo, Ben Phan, Jung-Hsien Chiang
Biomedical relation extraction often involves datasets with implicit constraints, where structural, syntactic, or semantic rules must be strictly preserved to maintain data integrity. Traditional data augmentation techniques struggle in these scenarios, as they risk violating domain-specific constraints. To address…
Brian Kenji Iwana, Seiichi Uchida, Friedhelm Schwenker
In recent times, deep artificial neural networks have achieved many successes in pattern recognition. Part of this success can be attributed to the reliance on big data to increase generalization. However, in the field of time series recognition, many datasets are often very small. One method of addressing this problem…
Hyunjung Lee, Utku Ozbulak, Homin Park, Stephen Depuydt + 2 more
'Wesley De Neve' 'Joris Vankerschaver'] Background Deep neural networks (DNNs) have the potential to revolutionize our understanding and treatment of genetic diseases. An inherent limitation of deep neural networks, however, is their high demand for data during training. To overcome this challenge, other fields, such…
Ferhat Ozgur Catak, Javed Ahmed, Kevser Sahinbas, Zahid Hussain Khand + 1 more
Due to advancements in malware competencies, cyber-attacks have been broadly observed in the digital world. Cyber-attacks can hit an organization hard by causing several damages such as data breach, financial loss, and reputation loss. Some of the most prominent examples of ransomware attacks in history are WannaCry…
Luka Lukač, Andrej Nerat, Damjan Strnad, Ivana Kolingerová + 2 more
This article presents a novel method for direct noise injection into geometric shapes described by eight-or four-directional Freeman chain codes. Noise is applied to randomly selected segments of a chain code sequence using a set of predefined actions. The design of alterations retains topological characteristics of…
Chinmayee Athalye, Rima Arnaout, Kathiravan Srinivasan
While domain-specific data augmentation can be useful in training neural networks for medical imaging tasks, such techniques have not been widely used to date. Our objective was to test whether domain-specific data augmentation is useful for medical imaging using a well-benchmarked task: view classification on fetal…
Weihao Li, Dan Jiang, Han Zhang, Kejing Xiao + 2 more
'Bilal Alatas'] The dialogue summarization is necessary for information retrieval, and the training of abstract dialogue summarization models heavily rely on large amounts of labeled data. However, manual summarization of long dialogue is labor-costing and time-consuming. To solve this problem, this article proposes a…
Haoyu Xiong, Xinchun Zhang, Leixin Yang, Yu Xiang + 2 more
'Arkaitz Zubiaga'] Test-time augmentation (TTA) is a well-established technique that involves aggregating transformed examples of test inputs during the inference stage. The goal is to enhance model performance and reduce the uncertainty of predictions. Despite its advantages of not requiring additional training or…
Man-Fai Wong, Shangxin Guo, Ching-Nam Hang, Siu-Wai Ho + 2 more
'Chee-Wei Tan' 'Lei Wang'] This paper provides a comprehensive review of the literature concerning the utilization of Natural Language Processing (NLP) techniques, with a particular focus on transformer-based large language models (LLMs) trained using Big Code, within the domain of AI-assisted programming tasks. LLMs…
Claudia Negri-Ribalta, Rémi Geraud-Stewart, Anastasia Sergeeva, Gabriele Lenzini
Introduction Artificial Intelligence (AI) is increasingly used as a helper to develop computing programs. While it can boost software development and improve coding proficiency, this practice offers no guarantee of security. On the contrary, recent research shows that some AI models produce software with…
R. Sakthivel, Ch. Vijayalakshmi, M. Vanitha, Kareem M. AboRas + 4 more
'Waleed Mohammed Abdelfattah' 'Yazeed Yasin Ghadi' 'Ch. Rami Reddy' 'Shonak Bansal'] Loss-less data compression becomes the need of the hour for effective data compression and computation in VLSI test vector generation and testing in addition to hardware AI/ML computations. Golomb code is one of the effective technique…