12 papers · ranked by Valyu relevance
Qihong Song, Haize Hu, Tebo Dai
Code search aims to search for code snippets from large codebase that are semantically related to natural query statements. Deep learning is a valuable method for solving code search tasks in which the quality of training data directly impacts the performance of deep-learning models. However, most existing…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Fang-Yi Su, Gia-Han Ngo, Ben Phan, Jung-Hsien Chiang
Biomedical relation extraction often involves datasets with implicit constraints, where structural, syntactic, or semantic rules must be strictly preserved to maintain data integrity. Traditional data augmentation techniques struggle in these scenarios, as they risk violating domain-specific constraints. To address…
Hyunjung Lee, Utku Ozbulak, Homin Park, Stephen Depuydt + 2 more
'Wesley De Neve' 'Joris Vankerschaver'] Background Deep neural networks (DNNs) have the potential to revolutionize our understanding and treatment of genetic diseases. An inherent limitation of deep neural networks, however, is their high demand for data during training. To overcome this challenge, other fields, such…
Luka Lukač, Andrej Nerat, Damjan Strnad, Ivana Kolingerová + 2 more
This article presents a novel method for direct noise injection into geometric shapes described by eight-or four-directional Freeman chain codes. Noise is applied to randomly selected segments of a chain code sequence using a set of predefined actions. The design of alterations retains topological characteristics of…
Haoyu Xiong, Xinchun Zhang, Leixin Yang, Yu Xiang + 2 more
'Arkaitz Zubiaga'] Test-time augmentation (TTA) is a well-established technique that involves aggregating transformed examples of test inputs during the inference stage. The goal is to enhance model performance and reduce the uncertainty of predictions. Despite its advantages of not requiring additional training or…
Chinmayee Athalye, Rima Arnaout, Kathiravan Srinivasan
While domain-specific data augmentation can be useful in training neural networks for medical imaging tasks, such techniques have not been widely used to date. Our objective was to test whether domain-specific data augmentation is useful for medical imaging using a well-benchmarked task: view classification on fetal…
Gang-Cheng Huang, Ko-Chin Chang, Tai-Hung Lai, Valderi R. Q. Leithardt
In this study, we propose a method for successfully evading antivirus detection by encoding malicious shellcode with fountain codes. The Meterpreter framework for Microsoft Windows 32-bit and 64-bit architectures was used to produce the shellcode used in this investigation. The experimental results proved that…
Weihao Li, Dan Jiang, Han Zhang, Kejing Xiao + 2 more
'Bilal Alatas'] The dialogue summarization is necessary for information retrieval, and the training of abstract dialogue summarization models heavily rely on large amounts of labeled data. However, manual summarization of long dialogue is labor-costing and time-consuming. To solve this problem, this article proposes a…
Man-Fai Wong, Shangxin Guo, Ching-Nam Hang, Siu-Wai Ho + 2 more
'Chee-Wei Tan' 'Lei Wang'] This paper provides a comprehensive review of the literature concerning the utilization of Natural Language Processing (NLP) techniques, with a particular focus on transformer-based large language models (LLMs) trained using Big Code, within the domain of AI-assisted programming tasks. LLMs…
Zhifeng Li, Yaqin Song, Runchen Li, Sen Gu + 4 more
'Longchao Cao' 'Qi Zhou'] Performing ultrasonic nondestructive testing experiments on insulators and then using machine learning algorithms to classify and identify the signals is an important way to achieve an intelligent diagnosis of insulators. However, in most cases, we can obtain only a limited number of data from…
R. Sakthivel, Ch. Vijayalakshmi, M. Vanitha, Kareem M. AboRas + 4 more
'Waleed Mohammed Abdelfattah' 'Yazeed Yasin Ghadi' 'Ch. Rami Reddy' 'Shonak Bansal'] Loss-less data compression becomes the need of the hour for effective data compression and computation in VLSI test vector generation and testing in addition to hardware AI/ML computations. Golomb code is one of the effective technique…