24 papers · ranked by Valyu relevance
Rhys Compton, Eibe Frank, Panos Patros, Abigail Koay
Automatic source code analysis in key areas of software engineering, such as code security, can benefit from Machine Learning (ML). However, many standard ML approaches require a numeric representation of data and cannot be applied directly to source code. Thus, to enable ML, we need to embed source code into numeric…
Peter Samoaa, Mehrdad Vasheghani Farahani, Antonio Longa, Philipp Leitner + 1 more
Tasks Authors: ['Peter Samoaa' 'Mehrdad Vasheghani Farahani' 'Antonio Longa' 'Philipp Leitner' 'Morteza Haghir Chehreghani'] Abstract—The landscape of deep learning has vastly expanded the frontiers of source code analysis, particularly through the utilization of structural representations such as Abstract Syntax Trees…
Wei Ma, Mengjie Zhao, Ezekiel Soremekun, Qiang Hu + 5 more
'Mike Papadakis' 'Maxime Cordy' 'Xiaofei Xie' 'Yves Le Traon'] Code embedding is a keystone in the application of machine learning on several Software Engineering (SE) tasks. To effectively support a plethora of SE tasks, the embedding needs to capture program syntax and semantics in a way that is generic. To this end…
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Nghi D. Q. Bui, Yijun Yu, Lingxiao Jiang
—Learning code representations has found many uses in software engineering, such as code classification, code search, code comment generation, and bug prediction. Although representations of code in tokens, syntax trees, dependency graphs, paths in trees, or the combinations of their variants have been proposed…
Abdullah Al Ishtiaq, Masum Hasan, Md. Mahim Anjum Haque, Kazi Sajeed Mehrab + 4 more
'Kazi Sajeed Mehrab' 'Tanveer Muttaqueen' 'Tahmid Hasan' 'Anindya Iqbal' 'Rifat Shahriyar'] Millions of repetitive code snippets are submitted to code repositories every day. To search from these large codebases using simple natural language queries would allow programmers to ideate, prototype, and develop easier and…
Nghi D. Q. Bui, Yijun Yu, Lingxiao Jiang
We propose Corder, a self-supervised contrastive learning framework for source code model. Corder is designed to alleviate the need of labeled data for code retrieval and code summarization tasks. The pre-trained model of Corder can be used in two ways: (1) it can produce vector representation of code which can be…
Muhammad Hammad, Önder Babur, Hamid Abdul Basit, Mark van den Brand + 1 more
'Yilun Shang'] Software developers frequently reuse source code from repositories as it saves development time and effort. Code clones (similar code fragments) accumulated in these repositories represent often repeated functionalities and are candidates for reuse in an exploratory or rapid development. To facilitate…
Andrea Gurioli, Maurizio Gabbrielli, Stefano Zacchiroli, Stefan Wagner
'Stefan Wagner'] Code stylometry is the application of stylometry techniques to determine the authorship of software source code snippets. It is used in the industry to address use cases like plagiarism detection, code audits, and code review assignments. Most works in the code stylometry literature use machine…
Qihong Song, Haize Hu, Tebo Dai
Code search aims to search for code snippets from large codebase that are semantically related to natural query statements. Deep learning is a valuable method for solving code search tasks in which the quality of training data directly impacts the performance of deep-learning models. However, most existing…
Binita Rajbanshi, Anuj Guruacharya
Emerging generative models for biology focus on DNA, non-coding RNA, or proteins, ignoring information hidden in mRNA. Additionally, in protein engineering and mRNA therapeutics the design of mRNA sequences is still a challenge, lacking a clear framework. Here, we introduce and rigorously evaluate two novel methods: a…
Marc Joiret, Marine Leclercq, Gaspard Lambrechts, Francesca Rapino + 3 more
'Pierre Close' 'Gilles Louppe' 'Liesbet Geris'] The genetic code is textbook scientific knowledge that was soundly established without resorting to Artificial Intelligence (AI). The goal of our study was to check whether a neural network could re-discover, on its own, the mapping links between codons and amino acids…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Tagir Akhmetshin, Arkadii Lin, Timur Madzhidov, Alexandre Varnek
Autoencoders represent a promising technique for the inverse quantitative structure-activity relationship (QSAR) task. However, undesirable bias, such as atom ordering, affects the neighbourhood behaviour of autoencoders’ latent space and, consequently, usage of the latent vectors as variables in machine-learning…
Logan Hallee, Nikolaos Rafailidis, Jason P. Gleghorn
Recent advancements in Protein Language Models (pLMs) have enabled high-throughput analysis of proteins through primary sequence alone. At the same time, newfound evidence illustrates that codon usage bias is remarkably predictive and can even change the final structure of a protein. Here, we explore these findings by…
Kristína Machová, Marián Mach, Michal Porezaný, Seongsoo Cho
This article focuses on the problem of detecting disinformation about COVID-19 in online discussions. As the Internet expands, so does the amount of content on it. In addition to content based on facts, a large amount of content is being manipulated, which negatively affects the whole society. This effect is currently…
Zhangyang Gao, Cheng Tan, Stan Z. Li
Protein structure tokenization has attracted increasing attention in both protein representation learning and generation. While recent work, like FoldToken2 and ESM3, has achieved good reconstruction performance, the compressoin ratio is still limited. In this work, we propose FoldToken3, a novel protein structure…
Jianhui Zeng, Zhiheng Qu, Bo Cai, Boris Ryabko
Source code summarization focuses on generating qualified natural language descriptions of a code snippet (e.g., functionality, usage and version). In an actual development environment, descriptions of the code are missing or not consistent with the code due to human factors, which makes it difficult for developers to…
Authors not listed
Compound similarity is fundamental to various cheminformatics analyses, particularly in the drug discovery industry, where the structure-activity principle is central to medicinal chemistry. Historically, binary fingerprints combined with Tanimoto and “Tanimoto-related metrics” (such as Dice, Sørensen–Dice, and…
Yanshuo Chen, Yuming Zhang, Joshua Li, Boxue Tian + 1 more
Codon optimization involves selecting synonymous codons to match host-specific preferences. It is critical for heterologous expression but remains challenging due to the combinatorial design space. Under long-term evolutionary selection, natural coding sequences are near-optimal compromises between translational…
Melissa Sanabria, Jonas Hirsch, Anna R. Poetsch
Large Language Models (LLMs) on natural language have achieved a level of performance that allows the generation of coherent and syntactically correct text. DNA sequence of genomes follows rules similar to natural language, but a distinguishing factor is the absence of a concept analogous to words. We established…
Jiayi Li, Hong-Sheng Lai, Litian Liang, Shiyi Du + 2 more
Codon sequence design is crucial for generating mRNA sequences with desired functional properties for tasks such as developing mRNA vaccines or gene editing therapies. Yet existing methods lack flexibility and controllability to adapt to various design objectives. We propose a novel machine learning-based framework…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…