20 papers · ranked by Valyu relevance
Li HaoChen, Chunyan Miao, Cyril Leung, Yanxian Huang + 3 more
'Hongyu Zhang' 'Yanlin Wang'] Code search, which aims at retrieving the most relevant code fragment for a given natural language query, is a common activity in software development practice. Recently, contrastive learning is widely used in code search research, where many data augmentation approaches for source code…
Zeming Dong, Qiang Hu, Yuejun Guo, Maxime Cordy + 3 more
'Yves Le Traon' 'Jianjun Zhao'] Abstract—Inspired by the great success of Deep Neural Networks (DNNs) in natural language processing (NLP), DNNs have been increasingly applied in source code analysis and attracted significant attention from the software engineering community. Due to its data-driven nature, a DNN model…
Joonghyuk Hahn, Hyeseon Ahn, Jungin Kim, Soohan Lim + 1 more
Time complexity is a theoretic measure to determine the amount of time the algorithm needs for its execution. In reality, developers write algorithms into code snippets within limited resources, making the calculation of a code's time complexity a fundamental task. However, determining the precise time complexity of a…
Qihong Song, Haize Hu, Tebo Dai
Code search aims to search for code snippets from large codebase that are semantically related to natural query statements. Deep learning is a valuable method for solving code search tasks in which the quality of training data directly impacts the performance of deep-learning models. However, most existing…
Yanlin Wang, Lianghong Guo, Ensheng Shi, Wenqing Chen + 7 more
'Wanjun Zhong' 'Menghan Wang' 'Hui Li' 'Hongyu Zhang' 'Ziyu Lyu' 'Zibin Zheng'] Abstract—Code search plays a crucial role in software development, enabling developers to retrieve and reuse code using natural language queries. While the performance of code search models improves with an increase in high-quality data…
Zeming Dong, Qiang Hu, Yuejun Guo, Zhenya Zhang + 4 more
'Mike Papadakis' 'Yves Le Traon' 'Jianjun Zhao'] The next era of program understanding is being propelled by the use of machine learning to solve software problems. Recent studies have shown surprising results of source code learning, which applies deep neural networks (DNNs) to various critical software tasks, e.g.…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Fang-Yi Su, Gia-Han Ngo, Ben Phan, Jung-Hsien Chiang
Biomedical relation extraction often involves datasets with implicit constraints, where structural, syntactic, or semantic rules must be strictly preserved to maintain data integrity. Traditional data augmentation techniques struggle in these scenarios, as they risk violating domain-specific constraints. To address…
Zezhou Yang, Sirong Chen, Cuiyun Gao, Zhenhao Li + 3 more
and Opportunities Authors: ['Zezhou Yang' 'Sirong Chen' 'Cuiyun Gao' 'Zhenhao Li' 'Xing Hu' 'Kui Liu' 'Xin Xia'] Code generation aims to automatically generate code snippets of specific programming language according to natural language descriptions. The continuous advancements in deep learning, particularly…
Hyunjung Lee, Utku Ozbulak, Homin Park, Stephen Depuydt + 2 more
'Wesley De Neve' 'Joris Vankerschaver'] Background Deep neural networks (DNNs) have the potential to revolutionize our understanding and treatment of genetic diseases. An inherent limitation of deep neural networks, however, is their high demand for data during training. To overcome this challenge, other fields, such…
Henning Otto Brinkhaus, Kohulan Rajan, Achim Zielesny, Christoph Steinbeck
The development of deep learning-based optical chemical structure recognition (OCSR) systems has led to a need for datasets of chemical structure depictions. The diversity of the features in the training data is an important factor for the generation of deep learning systems that generalise well and are not overfit to…
Authors not listed
Data augmentation can alleviate the limitations of small molecular datasets for generative deep learning, by ‘artificially inflating’ the number of instances available for training. SMILES enumeration – whereby multiple valid SMILES strings are used to represent the same molecules – has resulted particularly beneficial…
Haoyu Xiong, Xinchun Zhang, Leixin Yang, Yu Xiang + 2 more
'Arkaitz Zubiaga'] Test-time augmentation (TTA) is a well-established technique that involves aggregating transformed examples of test inputs during the inference stage. The goal is to enhance model performance and reduce the uncertainty of predictions. Despite its advantages of not requiring additional training or…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…
Akos Nyerges, Anush Chiappino-Pepe, Bogdan Budnik, Maximilien Baas-Thomas + 26 more
Engineering the genetic code of an organism provides the basis for (i) making any organism safely resistant to natural viruses and (ii) preventing genetic information flow into and out of genetically modified organisms while (iii) allowing the biosynthesis of genetically encoded unnatural polymers^1–4^. Achieving these…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Pawel M. Mordaka, Kitty Clouston, Jing Cui, Andre Holzer + 3 more
Genome scale engineering has enabled codon compression of the universal genetic code of up to three codons in E. coli, providing the means for genetic code expansion. To go much beyond this number, smaller and simpler genetic systems are needed to avoid significant technical challenges. Chloroplast genomes offer…
Jeff Guo, Philippe Schwaller
Sample efficiency is a fundamental challenge in de novo molecular design. Ideally, molecular generative models should learn to satisfy desired objectives under minimal oracle evaluations (computational prediction or wet-lab experiment). This problem becomes more apparent when using oracles that can provide increased…
Jeff Guo, Philippe Schwaller
Sample efficiency is a fundamental challenge in de novo molecular design. Ideally, molecular generative models should learn to satisfy desired objectives under minimal oracle evaluations (computational prediction or wet-lab experiment). This problem becomes more apparent when using oracles that can provide increased…
Yunran Chen, Jennifer M Groh, Surya T Tokdar
Understanding how neurons encode multiple simultaneous stimuli is a fundamental question in neuroscience. We have previously introduced a novel theory of stochastic encoding patterns wherein a neuron’s spiking activity dynamically switches among its constituent single-stimulus activity patterns when presented with…