23 papers · ranked by Valyu relevance
Authors not listed
Vulnerability code-bases often suffer from severe imbalance, limiting the effectiveness of Deep Learning-based vulnerability classifiers. Data Augmentation could help solve this by mitigating the scarcity of under-represented CWEs. In this context, we investigate LLM-based augmentation for vulnerable functions…
Li HaoChen, Chunyan Miao, Cyril Leung, Yanxian Huang + 3 more
'Hongyu Zhang' 'Yanlin Wang'] Code search, which aims at retrieving the most relevant code fragment for a given natural language query, is a common activity in software development practice. Recently, contrastive learning is widely used in code search research, where many data augmentation approaches for source code…
Joonghyuk Hahn, Hyeseon Ahn, Jungin Kim, Soohan Lim + 1 more
Time complexity is a theoretic measure to determine the amount of time the algorithm needs for its execution. In reality, developers write algorithms into code snippets within limited resources, making the calculation of a code's time complexity a fundamental task. However, determining the precise time complexity of a…
Qihong Song, Haize Hu, Tebo Dai
Code search aims to search for code snippets from large codebase that are semantically related to natural query statements. Deep learning is a valuable method for solving code search tasks in which the quality of training data directly impacts the performance of deep-learning models. However, most existing…
Yanlin Wang, Lianghong Guo, Ensheng Shi, Wenqing Chen + 7 more
'Wanjun Zhong' 'Menghan Wang' 'Hui Li' 'Hongyu Zhang' 'Ziyu Lyu' 'Zibin Zheng'] Abstract—Code search plays a crucial role in software development, enabling developers to retrieve and reuse code using natural language queries. While the performance of code search models improves with an increase in high-quality data…
Zeming Dong, Qiang Hu, Yuejun Guo, Maxime Cordy + 3 more
'Yves Le Traon' 'Jianjun Zhao'] Abstract—Inspired by the great success of Deep Neural Networks (DNNs) in natural language processing (NLP), DNNs have been increasingly applied in source code analysis and attracted significant attention from the software engineering community. Due to its data-driven nature, a DNN model…
Mehdi Bahrami, N. C. Shrikanth, Yuji Mizobuchi, Lei Liu + 3 more
'Masahiro Fukuyori' 'Weipeng Chen' 'Kazuki Munakata'] Code retrieval is allowing software engineers to search codes through a natural language query, which relies on both natural language processing and software engineering techniques. There have been several attempts on code retrieval from searching snippet codes to…
Zeming Dong, Qiang Hu, Yuejun Guo, Zhenya Zhang + 4 more
'Mike Papadakis' 'Yves Le Traon' 'Jianjun Zhao'] The next era of program understanding is being propelled by the use of machine learning to solve software problems. Recent studies have shown surprising results of source code learning, which applies deep neural networks (DNNs) to various critical software tasks, e.g.…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Fang-Yi Su, Gia-Han Ngo, Ben Phan, Jung-Hsien Chiang
Biomedical relation extraction often involves datasets with implicit constraints, where structural, syntactic, or semantic rules must be strictly preserved to maintain data integrity. Traditional data augmentation techniques struggle in these scenarios, as they risk violating domain-specific constraints. To address…
Brian Kenji Iwana, Seiichi Uchida, Friedhelm Schwenker
In recent times, deep artificial neural networks have achieved many successes in pattern recognition. Part of this success can be attributed to the reliance on big data to increase generalization. However, in the field of time series recognition, many datasets are often very small. One method of addressing this problem…
Paweł Błażej, Małgorzata Wnetrzak, Dorota Mackiewicz, Paweł Mackiewicz
Compounds including non-canonical amino acids or other artificially designed molecules can find a lot of applications in medicine, industry and biotechnology. They can be produced thanks to the modification or extension of the standard genetic code (SGC). Such peptides or proteins including the non-canonical amino…
Hyunjung Lee, Utku Ozbulak, Homin Park, Stephen Depuydt + 2 more
'Wesley De Neve' 'Joris Vankerschaver'] Background Deep neural networks (DNNs) have the potential to revolutionize our understanding and treatment of genetic diseases. An inherent limitation of deep neural networks, however, is their high demand for data during training. To overcome this challenge, other fields, such…
Yiyang Yu, Shivani Muthukumar, Peter K Koo
Deep neural networks (DNNs) have been widely applied to predict the molecular functions of regulatory regions in the non-coding genome. DNNs are data hungry and thus require many training examples to fit data well. However, functional genomics experiments typically generate limited amounts of data, constrained by the…
Nikita Janakarajan, Mara Graziani, Maria Rodriguez Martinez
Working with transcriptomic data is challenging in deep learning applications due to its high dimensionality and low patient numbers. Deep learning models tend to overfit this data and do not generalize well on out-of-distribution samples and new cohorts. Data augmentation strategies help alleviate this problem by…
Authors not listed
Data augmentation can alleviate the limitations of small molecular datasets for generative deep learning, by ‘artificially inflating’ the number of instances available for training. SMILES enumeration – whereby multiple valid SMILES strings are used to represent the same molecules – has resulted particularly beneficial…
Andrew G Duncan, Jennifer A Mitchell, Alan M Moses
Supervised deep learning is used to model the complex relationship between genomic sequence and regulatory function. Understanding how these models make predictions can provide biological insight into regulatory functions. Given the complexity of the sequence to regulatory function mapping (the cis-regulatory code), it…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
Esben Bjerrum, Tobias Rastemo, Ross Irwin, Christos Kannas + 1 more
Recent years have seen a large interest in using the Simplified Molecular Input Line Entry System (SMILES) chemical language as input for deep learning architectures solving chemical tasks. Many successful applications have been demonstrated within de novo molecular design, quantitative structure-activity relationship…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Pawel M. Mordaka, Kitty Clouston, Jing Cui, Andre Holzer + 3 more
Genome scale engineering has enabled codon compression of the universal genetic code of up to three codons in E. coli, providing the means for genetic code expansion. To go much beyond this number, smaller and simpler genetic systems are needed to avoid significant technical challenges. Chloroplast genomes offer…
Jeff Guo, Philippe Schwaller
Sample efficiency is a fundamental challenge in de novo molecular design. Ideally, molecular generative models should learn to satisfy desired objectives under minimal oracle evaluations (computational prediction or wet-lab experiment). This problem becomes more apparent when using oracles that can provide increased…
Jeff Guo, Philippe Schwaller
Sample efficiency is a fundamental challenge in de novo molecular design. Ideally, molecular generative models should learn to satisfy desired objectives under minimal oracle evaluations (computational prediction or wet-lab experiment). This problem becomes more apparent when using oracles that can provide increased…