20 papers · ranked by Valyu relevance
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Zimin Chen, Martin Monperrus
Natural language processing has improved tremendously after the success of word embedding techniques such as word2vec. Recently, the same idea has been applied on source code with encouraging results. In this survey, we aim to collect and discuss the usage of word embedding techniques on programs and source code. The…
Maryam Vahdat Pour, Zhuo Li, Lei Ma, Hadi Hemmati
—Over the past few years, deep neural networks (DNNs) have been continuously expanding their real-world applications for source code processing tasks across the software engineering domain, e.g., clone detection, code search, comment generation. Although quite a few recent works have been performed on testing of DNNs…
Abdullah Al Ishtiaq, Masum Hasan, Md. Mahim Anjum Haque, Kazi Sajeed Mehrab + 4 more
'Kazi Sajeed Mehrab' 'Tanveer Muttaqueen' 'Tahmid Hasan' 'Anindya Iqbal' 'Rifat Shahriyar'] Millions of repetitive code snippets are submitted to code repositories every day. To search from these large codebases using simple natural language queries would allow programmers to ideate, prototype, and develop easier and…
Muhammad Usama, Ulas Yaman, Patricia Krawczak
The paper gives a detailed review of the approaches adopted for embedding information into/onto additively manufactured parts. The primary purpose of this paper is to review all the techniques adopted for embedding information, highlight notable trends and improvements in these works, and provide design and…
Yuhao Jia, Zhicheng Yu, Zhen Hong, Asadullah Shaikh
Binary code similarity detection plays a crucial role in various applications within binary security, including vulnerability detection, malicious software analysis, etc. However, existing methods suffer from limited differentiation in binary embedding representations across different compilation environments, lacking…
Xueyan Tang, Yuying Du, Alan Lai, Ze Zhang + 1 more
This paper aims to explore the application of deep learning in smart contract vulnerabilities detection. Smart contracts are an essential part of blockchain technology and are crucial for developing decentralized applications. However, smart contract vulnerabilities can cause financial losses and system crashes. Static…
Ruibo Shi, Lili Tao, Rohan Saphal, Fran Silavong + 1 more
We present CV4Code, a compact and effective computer vision method for sourcecode understanding. Our method leverages the contextual and the structural information available from the code snippet by treating each snippet as a two-dimensional image, which naturally encodes the context and retains the underlying…
Bing Xia, Jianmin Pang, Xin Zhou, Zheng Shan + 2 more
'Feng Yue'] Binary code similarity analysis is widely used in the field of vulnerability search where source code may not be available to detect whether two binary functions are similar or not. Based on deep learning and natural processing techniques, several approaches have been proposed to perform cross-platform…
Rhys Compton, Eibe Frank, Panos Patros, Abigail Koay
Automatic source code analysis in key areas of software engineering, such as code security, can benefit from Machine Learning (ML). However, many standard ML approaches require a numeric representation of data and cannot be applied directly to source code. Thus, to enable ML, we need to embed source code into numeric…
David Haughton, Félix Balado
Background In recent times, the application of deoxyribonucleic acid (DNA) has diversified with the emergence of fields such as DNA computing and DNA data embedding. DNA data embedding, also known as DNA watermarking or DNA steganography, aims to develop robust algorithms for encoding non-genetic information in DNA.…
Niklas Brunn, Sonia Maria Krißmer, Maximilian Frosch, Markus Frick + 2 more
The single-cell literature catalogs cell states as validated marker-gene programs — a sparse, compositional prior. Conventional embedding methods do not leverage this prior and learn cell-state structure de novo from the expression matrix, producing dense dimensions needing post-hoc interpretation and batch correction.…
Rachel D. Melamed
The electronic health record is a rising resource for quantifying medical practice and discovering adverse effects of drugs. One of the challenges of applying these methods to health care data is the high dimensionality of the health record. Methods to discover effects of drugs in health data must account for tens of…
Zekun Cao, Zhaoxia Yin, Honghe Hu, Xiangping Gao + 1 more
Aiming to embed large amount of data while minimize the sum of costs of all changed pixels, a novel high capacity data hiding scheme based on (7, 4) Hamming code is realized by a family of algorithms. Firstly, n (n = 1, 2, 3) cover pixels are assigned to one set according to the payload. Then, 128 binary strings of…
Authors not listed
High-level quantum mechanical (QM) simulations provide accurate electronic information of chemical systems but scale unfavourably with system size, making calculations of applied systems challenging. Hierarchical quantum mechanics in quantum mechanics embedding (QM/QM) addresses this issue by localising the highly…
Melissa Franch, Elizabeth A. Mickiewicz, James L. Belanger, Assia Chericoni + 10 more
As we listen to speech, our brains actively compute the meaning of individual words. Inspired by the success of large language models (LLMs), we hypothesized that the brain employs vectorial coding principles, such that meaning is reflected in distributed activity of single neurons. We recorded responses of hundreds of…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Zhe Liu, Yihang Bao, Wenhao Li, Weihao Li + 1 more
Non-coding single nucleotide polymorphisms (SNPs) are critical drivers of gene regulation and disease susceptibility, yet predicting their functional impact remains a challenging task. A variety of methods exist for encoding non-coding SNPs, such as direct base encoding or using pre-trained models to obtain embeddings.…
Y. Liu, J. Kim, C. Wilson, M. Bedny
Despite the importance of programming to modern society, the cognitive and neural bases of code comprehension are largely unknown. Programming languages might ‘recycle’ neurocognitive mechanisms originally used for natural languages. Alternatively, comprehension of code could depend on fronto-parietal networks shared…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…