22 papers · ranked by Valyu relevance
Yutao Xie, Jiayi Lin, Hande Dong, Lei Zhang + 1 more
Code writing is repetitive and predictable, inspiring us to develop various code intelligence techniques. This survey focuses on code search, that is, to retrieve code that matches a given natural language query by effectively capturing the semantic similarity between the query and code. Deep learning, being able to…
Peter Samoaa, Mehrdad Vasheghani Farahani, Antonio Longa, Philipp Leitner + 1 more
Tasks Authors: ['Peter Samoaa' 'Mehrdad Vasheghani Farahani' 'Antonio Longa' 'Philipp Leitner' 'Morteza Haghir Chehreghani'] Abstract—The landscape of deep learning has vastly expanded the frontiers of source code analysis, particularly through the utilization of structural representations such as Abstract Syntax Trees…
Sairamvinay Vijayaraghavan, Jinxiao Song, David A. Tomassi, Siddhartha Punj + 1 more
Classification of text is an important field of research and a core task in Natural Language Processing (NLP). It spans many different domains from determining "fake" news, finding spam emails, and language detection. A common problem in software development is generating the appropriate code snippet for a task. There…
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Muhammad Hammad, Önder Babur, Hamid Abdul Basit, Mark van den Brand + 1 more
'Yilun Shang'] Software developers frequently reuse source code from repositories as it saves development time and effort. Code clones (similar code fragments) accumulated in these repositories represent often repeated functionalities and are candidates for reuse in an exploratory or rapid development. To facilitate…
Md Rafiqul Islam Rabin, Mohammad Amin Alipour
—There are several approaches for encoding source code in the input vectors of neural models. These approaches attempt to include various syntactic and semantic features of input programs in their encoding. In this paper, we investigate CODE2SNAPSHOT, a novel representation of the source code that is based on the…
Zhiwei Xu, Min Zhou, Xibin Zhao, Yang Chen + 2 more
Code representations (a.k.a., embeddings) is of great importance in deep learning-based software engineering techniques. A highquality representation model can significantly improve the performance of many downstream tasks, such as code search [13, 23, 42], code clone detection [21, 48, 54], and bug localization [27].…
Lili Liu, Zhen Li, Yu Wen, Penglong Chen + 1 more
Software vulnerabilities have led to system attacks and data leakage incidents, and software vulnerabilities have gradually attracted attention. Vulnerability detection had become an important research direction. In recent years, Deep Learning (DL)-based methods had been applied to vulnerability detection. The DL-based…
Andrea Gurioli, Maurizio Gabbrielli, Stefano Zacchiroli, Stefan Wagner
'Stefan Wagner'] Code stylometry is the application of stylometry techniques to determine the authorship of software source code snippets. It is used in the industry to address use cases like plagiarism detection, code audits, and code review assignments. Most works in the code stylometry literature use machine…
Binita Rajbanshi, Anuj Guruacharya
Emerging generative models for biology focus on DNA, non-coding RNA, or proteins, ignoring information hidden in mRNA. Additionally, in protein engineering and mRNA therapeutics the design of mRNA sequences is still a challenge, lacking a clear framework. Here, we introduce and rigorously evaluate two novel methods: a…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Tagir Akhmetshin, Arkadii Lin, Timur Madzhidov, Alexandre Varnek
Autoencoders represent a promising technique for the inverse quantitative structure-activity relationship (QSAR) task. However, undesirable bias, such as atom ordering, affects the neighbourhood behaviour of autoencoders’ latent space and, consequently, usage of the latent vectors as variables in machine-learning…
Logan Hallee, Nikolaos Rafailidis, Jason P. Gleghorn
Recent advancements in Protein Language Models (pLMs) have enabled high-throughput analysis of proteins through primary sequence alone. At the same time, newfound evidence illustrates that codon usage bias is remarkably predictive and can even change the final structure of a protein. Here, we explore these findings by…
Kristína Machová, Marián Mach, Michal Porezaný, Seongsoo Cho
This article focuses on the problem of detecting disinformation about COVID-19 in online discussions. As the Internet expands, so does the amount of content on it. In addition to content based on facts, a large amount of content is being manipulated, which negatively affects the whole society. This effect is currently…
Zhangyang Gao, Cheng Tan, Stan Z. Li
Protein structure tokenization has attracted increasing attention in both protein representation learning and generation. While recent work, like FoldToken2 and ESM3, has achieved good reconstruction performance, the compressoin ratio is still limited. In this work, we propose FoldToken3, a novel protein structure…
Jianhui Zeng, Zhiheng Qu, Bo Cai, Boris Ryabko
Source code summarization focuses on generating qualified natural language descriptions of a code snippet (e.g., functionality, usage and version). In an actual development environment, descriptions of the code are missing or not consistent with the code due to human factors, which makes it difficult for developers to…
Authors not listed
Compound similarity is fundamental to various cheminformatics analyses, particularly in the drug discovery industry, where the structure-activity principle is central to medicinal chemistry. Historically, binary fingerprints combined with Tanimoto and “Tanimoto-related metrics” (such as Dice, Sørensen–Dice, and…
Yanshuo Chen, Yuming Zhang, Joshua Li, Boxue Tian + 1 more
Codon optimization involves selecting synonymous codons to match host-specific preferences. It is critical for heterologous expression but remains challenging due to the combinatorial design space. Under long-term evolutionary selection, natural coding sequences are near-optimal compromises between translational…
Melissa Sanabria, Jonas Hirsch, Anna R. Poetsch
Large Language Models (LLMs) on natural language have achieved a level of performance that allows the generation of coherent and syntactically correct text. DNA sequence of genomes follows rules similar to natural language, but a distinguishing factor is the absence of a concept analogous to words. We established…
Jiayi Li, Hong-Sheng Lai, Litian Liang, Shiyi Du + 2 more
Codon sequence design is crucial for generating mRNA sequences with desired functional properties for tasks such as developing mRNA vaccines or gene editing therapies. Yet existing methods lack flexibility and controllability to adapt to various design objectives. We propose a novel machine learning-based framework…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…