17 papers · ranked by Valyu relevance
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis + 1 more
'Marc Brockschmidt'] To enable evaluation of progress on code search, we are releasing the CodeSearchNet Corpus and are presenting the CodeSearch-Net Challenge, which consists of 99 natural language queries with about 4k expert relevance annotations of likely results from Code-SearchNet Corpus. The corpus contains…
Chen Wu, Ming Yan
Semantic code search is the task of retrieving relevant code snippet given a natural language query. Different from typical information retrieval tasks, code search requires to bridge the semantic gap between the programming language and natural language, for better describing intrinsic concepts and semantics.…
Yutao Xie, Jiayi Lin, Hande Dong, Lei Zhang + 1 more
Code writing is repetitive and predictable, inspiring us to develop various code intelligence techniques. This survey focuses on code search, that is, to retrieve code that matches a given natural language query by effectively capturing the semantic similarity between the query and code. Deep learning, being able to…
Andor Diera, Abdelhalim Hafedh Dahou, Lukas Galke, Fabian Karl + 2 more
'Florian Sihler' 'Ansgar Scherp'] Language models can serve as a valuable tool for software developers to increase productivity. Large generative models can be used for code generation and code completion, while smaller encoder-only models are capable of performing code search tasks using natural language queries.…
Ivan Sedykh, Dmitry Abulkhanov, Nikita Sorokin, Sergey Nikolenko + 1 more
'Valentin Malykh'] Code search is an important task that has seen many developments in recent years. However, previous attempts have mostly considered the problem of searching for code by a text query. We argue that using a code snippet (and possibly an associated traceback) as a query and looking for answers with…
Qihong Song, Haize Hu, Tebo Dai
Code search aims to search for code snippets from large codebase that are semantically related to natural query statements. Deep learning is a valuable method for solving code search tasks in which the quality of training data directly impacts the performance of deep-learning models. However, most existing…
Shushan Arakelyan, Anna Hakhverdyan, Miltiadis Allamanis, Christophe Hauser + 2 more
'Christophe Hauser' 'Luis García' 'Xiang Ren'] Semantic code search is the task of retrieving a code snippet given a textual description of its functionality. Recent work has been focused on using similarity metrics between neural embeddings of text and code. However, current language models are known to struggle with…
Muhammad Hammad, Önder Babur, Hamid Abdul Basit, Mark van den Brand + 1 more
'Yilun Shang'] Software developers frequently reuse source code from repositories as it saves development time and effort. Code clones (similar code fragments) accumulated in these repositories represent often repeated functionalities and are candidates for reuse in an exploratory or rapid development. To facilitate…
Anastasia Drozdova, Ekaterina Trofimova, Polina Guseva, Anna Scherbakova + 2 more
The use of program code as a data source is increasingly expanding among data scientists. The purpose of the usage varies from the semantic classification of code to the automatic generation of programs. However, the machine learning model application is somewhat limited without annotating the code snippets. To address…
Nazia Bibi, Tauseef Rana, Ayesha Maqbool, Farkhanda Afzal + 3 more
'Ali Akgül' 'Manuel De la Sen' 'Rebeca P. Díaz\xa0Redondo'] The development of robotic applications necessitates the availability of useful, adaptable, and accessible programming frameworks. Robotic, IoT, and sensor-based systems open up new possibilities for the development of innovative applications, taking advantage…
Zhengyu Zhao, Yuanyuan Lu, Yijie Tong, Xin Chen + 1 more
Discriminative traits are important in biodiversity and macroevolution, but extracting and representing these features from huge natural history collections using traditional methods can be challenging and time-consuming. To fully utilize the collections and their associated metadata, it is urgent now to increase the…
Evangelos Vlachos, Elliot Lefkowitz
Background In order to designate the various concepts of taxa in biology, evolution and paleontology, scientists have developed various rules on how to create unique names for taxa. Different Codes of Nomenclature have been developed for animals, plants, fungi, bacteria etc., with standard sets of Rules that govern the…
Yiwen Zhang, Wei Liu, Fazhong Jiang, Jiquan Ma + 4 more
Large Language Models of the Transformer architecture display great promise in automated code error detection based on their strength in processing sequential data. Nevertheless, their efficacy could be further improved by addressing the inherent weakness in handling structural code dependencies. In response to this…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Marc-Antoine Jacques, Maciej Dobrzyński, Paolo Armando Gagliardi, Raphael Sznitman + 1 more
Fluorescent biosensors routinely yield thousands of single-cell, heterogeneous, multi-dimensional signaling trajectories that are difficult to mine for relevant information. We present CODEX, an approach based on artificial neural networks to guide exploration of time-series datasets and to identify motifs in dynamic…
Bohdan B. Khomtchouk, Kasra A. Vand, Thor Wahlestedt, Kelly Khomtchouk + 2 more
We propose a search engine and file retrieval system for all bioinformatics databases worldwide. PubData searches biomedical data in a user-friendly fashion similar to how PubMed searches biomedical literature. PubData is built on novel network programming, natural language processing, and artificial intelligence…
Patrick Wu, Aliya Gifford, Xiangrui Meng, Xue Li + 7 more
Many studies of Electronic Health Record (EHR) data utilize custom-developed aggregations of billing codes enabling clinical and genetic research, including phenome-wide association studies (PheWAS). One such grouping is the phecode system, originally developed for PheWAS. Phecodes were built upon the International…