23 papers · ranked by Valyu relevance
José Cambronero, Hongyu Li, Seohyun Kim, Koushik Sen + 1 more
'Satish Chandra'] There have been multiple recent proposals on using deep neural networks for code search using natural language. Common across these proposals is the idea of embedding code and natural language queries into real vectors and then using vector distance to approximate semantic correlation between code and…
Zimin Chen, Martin Monperrus
Natural language processing has improved tremendously after the success of word embedding techniques such as word2vec. Recently, the same idea has been applied on source code with encouraging results. In this survey, we aim to collect and discuss the usage of word embedding techniques on programs and source code. The…
Zixiang Xian, Rubing Huang, Dave Towey, Chunrong Fang + 1 more
—Artificial intelligence (AI) has revolutionized software engineering (SE) by enhancing software development efficiency. The advent of pre-trained models (PTMs) leveraging transfer learning has significantly advanced AI for SE. However, existing PTMs that operate on individual code tokens suffer from several…
Abdullah Al Ishtiaq, Masum Hasan, Md. Mahim Anjum Haque, Kazi Sajeed Mehrab + 4 more
'Kazi Sajeed Mehrab' 'Tanveer Muttaqueen' 'Tahmid Hasan' 'Anindya Iqbal' 'Rifat Shahriyar'] Millions of repetitive code snippets are submitted to code repositories every day. To search from these large codebases using simple natural language queries would allow programmers to ideate, prototype, and develop easier and…
Wenchao Gu, Zongjie Li, Cuiyun Gao, Chaozheng Wang + 3 more
'Zenglin Xu' 'Michael R. Lyu'] Code retrieval is a common practice for programmers to reuse existing code snippets in the opensource repositories. Given a user query (i.e., a natural language description), code retrieval aims at searching the most relevant ones from a set of code snippets. The main challenge of…
Yuhao Jia, Zhicheng Yu, Zhen Hong, Asadullah Shaikh
Binary code similarity detection plays a crucial role in various applications within binary security, including vulnerability detection, malicious software analysis, etc. However, existing methods suffer from limited differentiation in binary embedding representations across different compilation environments, lacking…
Binita Rajbanshi, Anuj Guruacharya
Emerging generative models for biology focus on DNA, non-coding RNA, or proteins, ignoring information hidden in mRNA. Additionally, in protein engineering and mRNA therapeutics the design of mRNA sequences is still a challenge, lacking a clear framework. Here, we introduce and rigorously evaluate two novel methods: a…
Nghi D. Q. Bui, Yijun Yu, Lingxiao Jiang
We propose Corder, a self-supervised contrastive learning framework for source code model. Corder is designed to alleviate the need of labeled data for code retrieval and code summarization tasks. The pre-trained model of Corder can be used in two ways: (1) it can produce vector representation of code which can be…
Pawel Pratyush, Callen Carrier, Suresh Pokharel, Hamid D Ismail + 3 more
The highly degenerate nature of codons leads to a surjective mapping from multiple codons to a single amino acid, with most amino acids being encoded by up to six different codons (see [btaf124-F1]). This indicates that a sequence represented at the codon level might contain as much, if not more, information than the…
Rachel D. Melamed
The electronic health record is a rising resource for quantifying medical practice and discovering adverse effects of drugs. One of the challenges of applying these methods to health care data is the high dimensionality of the health record. Methods to discover effects of drugs in health data must account for tens of…
Andrea Gurioli, Maurizio Gabbrielli, Stefano Zacchiroli, Stefan Wagner
'Stefan Wagner'] Code stylometry is the application of stylometry techniques to determine the authorship of software source code snippets. It is used in the industry to address use cases like plagiarism detection, code audits, and code review assignments. Most works in the code stylometry literature use machine…
Ashis Kumar Chanda, Tian Bai, Ziyu Yang, Slobodan Vucetic
Background Health providers create Electronic Health Records (EHRs) to describe the conditions and procedures used to treat their patients. Medical notes entered by medical staff in the form of free text are a particularly insightful component of EHRs. There is a great interest in applying machine learning tools on…
Zhao Shuai, Diao Xiaolin, Yuan Jing, Huo Yanni + 3 more
'Wang Yuxin' 'Zhao Wei'] Background Automated ICD coding on medical texts via machine learning has been a hot topic. Related studies from medical field heavily relies on conventional bag-of-words (BoW) as the feature extraction method, and do not commonly use more complicated methods, such as word2vec (W2V) and large…
Gunther Eysenbach, Yanshan Wang, David Mendes, Klerisson Paixao + 9 more
'Sujay Kakarmath' 'Chin Lin' 'Yu-Sheng Lou' 'Dung-Jang Tsai' 'Chia-Cheng Lee' 'Chia-Jung Hsu' 'Ding-Chung Wu' 'Mei-Chuen Wang' 'Wen-Hui Fang'] Background Most current state-of-the-art models for searching the International Classification of Diseases, Tenth Revision Clinical Modification (ICD-10-CM) codes use word…
Beiji Lu
Synonymous codons encode the same amino acid yet are used non-randomly across genomes, a phenomenon with well-documented functional consequences for translation efficiency and mRNA stability. Whether the information embedded in synonymous codon choice is recoverable from the internal representations of in-dependently…
Authors not listed
Compound similarity is fundamental to various cheminformatics analyses, particularly in the drug discovery industry, where the structure-activity principle is central to medicinal chemistry. Historically, binary fingerprints combined with Tanimoto and “Tanimoto-related metrics” (such as Dice, Sørensen–Dice, and…
Authors not listed
Recent years have seen a growing interest in machine learning approaches for chemical tasks. The best existing methods focus on building base models that combine molecular graphs (“2D structures”) with atomic coordinates in 3D to predict molecular properties, typically through pre-training followed by fine-tuning on…
Kevin Durrheim, Maria Schuld, Martin Mafunda, Sindisiwe Mazibuko
Word embeddings provide quantitative representations of word semantics and the associations between word meanings in text data, including in large repositories in media and social media archives. This article introduces social psychologists to word embedding research via a consideration of bias analysis, a topic of…
Amer El-Samman, Stijn De Baerdemacker
In deep learning methods, especially in the context of chemistry, there is an increasing urgency to uncover the hidden learning mechanisms often dubbed as ``black box." In this work, we show that graph models built on computational chemical data behave similar to natural language processing (NLP) models built on text…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Authors not listed
Molecular structure elucidation is a crucial but fundamentally challenging step in the characterization of materials given the large number of possible structures. Here, we introduce Spectro, an innovative multi-modal approach for molecular elucidation that combines $\CNMR$ and $\HNMR$ NMR data with IR. Spectro…
Y. Liu, J. Kim, C. Wilson, M. Bedny
Despite the importance of programming to modern society, the cognitive and neural bases of code comprehension are largely unknown. Programming languages might ‘recycle’ neurocognitive mechanisms originally used for natural languages. Alternatively, comprehension of code could depend on fronto-parietal networks shared…
Authors not listed
Real-world datasets in chemical engineering and bioengineering processes--such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials--can often be unlabelled or disorganized, rendering the training of existing supervised learning models ineffective at learning the…