20 papers · ranked by Valyu relevance
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Muhammad Usama, Ulas Yaman, Patricia Krawczak
The paper gives a detailed review of the approaches adopted for embedding information into/onto additively manufactured parts. The primary purpose of this paper is to review all the techniques adopted for embedding information, highlight notable trends and improvements in these works, and provide design and…
Ruibo Shi, Lili Tao, Rohan Saphal, Fran Silavong + 1 more
We present CV4Code, a compact and effective computer vision method for sourcecode understanding. Our method leverages the contextual and the structural information available from the code snippet by treating each snippet as a two-dimensional image, which naturally encodes the context and retains the underlying…
Ankit Kulshrestha, Vishwas Lele
There has been a steadily growing interest in development of novel methods to learn a representation of a given input data and subsequently using them for several downstream tasks. The field of natural language processing has seen a significant improvement in different tasks by incorporating pretrained embeddings into…
Saiteja Utpala, Alex Gu, Pin Yu Chen
Recently, code language models have achieved notable advancements in addressing a diverse array of essential code comprehension and generation tasks. Yet, the field lacks a comprehensive deep dive and understanding of the code embeddings of multilingual code models. In this paper, we present a comprehensive study on…
Yu Zhao, Lina Gong, Haoxiang Zhang, Yaoshen Yu + 1 more
Pre-trained language models have demonstrated powerful capabilities in the field of natural language processing (NLP). Recently, code pretrained model (PTM), which draw from the experiences of the NLP field, have also achieved state-of-the-art results in many software engineering (SE) downstream tasks. These code PTMs…
Chuan Hong, Everett Rush, Molei Liu, Doudou Zhou + 19 more
'Aaron Sonabend' 'Victor M. Castro' 'Petra Schubert' 'Vidul A. Panickan' 'Tianrun Cai' 'Lauren Costa' 'Zeling He' 'Nicholas Link' 'Ronald Hauser' 'J. Michael Gaziano' 'Shawn N. Murphy' 'George Ostrouchov' 'Yuk-Lam Ho' 'Edmon Begoli' 'Junwei Lu' 'Kelly Cho' 'Katherine P. Liao' 'Tianxi Cai' ''] The increasing…
Zhiwei Xu, Min Zhou, Xibin Zhao, Yang Chen + 2 more
Code representations (a.k.a., embeddings) is of great importance in deep learning-based software engineering techniques. A highquality representation model can significantly improve the performance of many downstream tasks, such as code search [13, 23, 42], code clone detection [21, 48, 54], and bug localization [27].…
Yuhao Jia, Zhicheng Yu, Zhen Hong, Asadullah Shaikh
Binary code similarity detection plays a crucial role in various applications within binary security, including vulnerability detection, malicious software analysis, etc. However, existing methods suffer from limited differentiation in binary embedding representations across different compilation environments, lacking…
Xueyan Tang, Yuying Du, Alan Lai, Ze Zhang + 1 more
This paper aims to explore the application of deep learning in smart contract vulnerabilities detection. Smart contracts are an essential part of blockchain technology and are crucial for developing decentralized applications. However, smart contract vulnerabilities can cause financial losses and system crashes. Static…
Bing Xia, Jianmin Pang, Xin Zhou, Zheng Shan + 2 more
'Feng Yue'] Binary code similarity analysis is widely used in the field of vulnerability search where source code may not be available to detect whether two binary functions are similar or not. Based on deep learning and natural processing techniques, several approaches have been proposed to perform cross-platform…
Khaled El Emam, Bradley Malin, Shirin Sarejloo, Muhammad Shahzad Aslam + 4 more
'Muhammad Shahzad Aslam' 'Sebastian Daniel Boie' 'Wei Zhang' 'Edgar Steiger' 'Lars Eric Kroll'] Background In health care, diagnosis codes in claims data and electronic health records (EHRs) play an important role in data-driven decision making. Any analysis that uses a patient’s diagnosis codes to predict future…
Authors not listed
High-level quantum mechanical (QM) simulations provide accurate electronic information of chemical systems but scale unfavourably with system size, making calculations of applied systems challenging. Hierarchical quantum mechanics in quantum mechanics embedding (QM/QM) addresses this issue by localising the highly…
Zhenhao Li, Hang Lei, Zhichao Ma, Fengyun Zhang + 4 more
'Yongpan Sheng' 'Hao Wang' 'Junyang Chen'] The code of industrial management software typically features few system API calls and a high number of customized variables and structures. This makes the similarity of such codes difficult to compute using text features or traditional neural network methods. In this paper…
Yun-Fei Liu, Marina Bedny
Programming is a cornerstone of modern society, yet its cognitive and neural basis remains poorly understood. In this study, we test the hypothesis that programming “recycles” pre-existing neural mechanisms and representations in fronto-parietal reasoning networks. Using fMRI, we scanned programming-naïve…
Logan Hallee, Nikolaos Rafailidis, Jason P. Gleghorn
Recent advancements in Protein Language Models (pLMs) have enabled high-throughput analysis of proteins through primary sequence alone. At the same time, newfound evidence illustrates that codon usage bias is remarkably predictive and can even change the final structure of a protein. Here, we explore these findings by…
Beiji Lu
Synonymous codons encode the same amino acid yet are used non-randomly across genomes, a phenomenon with well-documented functional consequences for translation efficiency and mRNA stability. Whether the information embedded in synonymous codon choice is recoverable from the internal representations of in-dependently…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Zhe Liu, Yihang Bao, Wenhao Li, Weihao Li + 1 more
Non-coding single nucleotide polymorphisms (SNPs) are critical drivers of gene regulation and disease susceptibility, yet predicting their functional impact remains a challenging task. A variety of methods exist for encoding non-coding SNPs, such as direct base encoding or using pre-trained models to obtain embeddings.…
Yann Spöri, Jean-François Flot
Haxe is a general purpose, object-oriented programming language supporting syntactic macros. The Haxe compiler is well known for its ability to translate the source code of Haxe programs into the source code of a variety of other programming languages including Java, C++, JavaScript and Python. Although Haxe is…