18 papers · ranked by Valyu relevance
Eduardo Guerra, Everaldo Mariano Gomes, Jeferson Ferreira, Igor Wiese + 3 more
'Igor Wiese' 'Phyllipe Lima' 'Marco Aurélio Gerosa' 'Paulo Meirelles'] Context: Code annotations have gained widespread popularity in programming languages, offering developers the ability to attach metadata to code elements to define custom behaviors. Many modern frameworks and APIs use annotations to keep integration…
Edward Misback, Erik Vank, Zachary Tatlock, Steven Tanimoto
When annotation data is stored along with code in the repository, teams immediately gain shared context, and annotations naturally follow version control history. However, such sharing also means that annotations become part of the codebase's footprint and could become a source of merge conflicts. This decision also…
Phyllipe Lima, Jorge Melegati, Everaldo Gomes, Nathalya Stefhany Pereira + 2 more
'Nathalya Stefhany Pereira' 'Eduardo Guerra' 'Paulo Meirelles'] Context: Code annotations is a widely used feature in Java systems to configure custom metadata on programming elements. Their increasing presence creates the need for approaches to assess and comprehend their usage and distribution. In this context…
Anastasia Drozdova, Ekaterina Trofimova, Polina Guseva, Anna Scherbakova + 2 more
The use of program code as a data source is increasingly expanding among data scientists. The purpose of the usage varies from the semantic classification of code to the automatic generation of programs. However, the machine learning model application is somewhat limited without annotating the code snippets. To address…
Jian Wang, Tao Lin, Rongsen Zhao, Huiling Zhao + 1 more
The deep natural language translation models have been used for automatic code error correction and have demonstrated outstanding potential. However, a large and accurately annotated training dataset is essential for these models to perform well. The key to improving the performance of these models lies in…
Amber Horvath, Michael Xieyang Liu, River Hendriksen, Connor Shannon + 4 more
'Emma Paterson' 'Kazi Jawad' 'Andrew Macvean' 'Brad A. Myers'] Modern software development requires developers to find and effectively utilize new APIs and their documentation, but documentation has many well-known issues. Despite this, developers eventually overcome these issues but have no way of sharing what they…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Paul C. Attie, Anas Obeidat, Nathaniel Oh, Ian Yelle
We present the Code Documentation and Analysis Tool (CoDAT). CoDAT is a tool designed to maintain consistency between the various levels of code documentation, e.g. if a line in a code sketch is changed, the comment that documents the corresponding code is also changed. That is, comments are linked and updated so as to…
M. Siepel, G.T.N. Burger, Q.J.M. Voorham, R. Cornet + 2 more
Palga Foundation is responsible for indexing Dutch pathology data across the Netherlands, which relies on annotations of pathology reports. These annotations, derived from the conclusion text, consist of codes from the Palga thesaurus, serving patient care and scientific research. However, manual annotation by…
Siwen Wu, Sisi Yuan, Wei-ming Li, Zhengchang Su
With the development of sequencing technology, genome assembly becomes more and more easy for individual labs. Many high-quality assemblies for various species have been released in NCBI these years, however, most of them lack annotations of protein-coding genes and pseudogenes due to the sophistication of the process.…
Nicolaï Hoffmann, Aurore Besson, Edouard Cadieu, Matthias Lorthiois + 6 more
With the advent of complete genome assemblies, genome annotation has become essential for the functional interpretation of genomic data. Long-read RNA sequencing (LR-RNAseq) technologies have significantly improved transcriptome annotation by enabling full-length transcript reconstruction for both coding and non-coding…
Lei Shi, Min Dai, Yongbo Zhang, Song Wu + 2 more
Single-cell omics and spatial omics technologies are nowadays widely used in biological and medical research. In both single-cell and spatial omics data analysis, accurate cell type annotation is a key step for downstream analysis and scientific discoveries. However, high-quality cell annotation usually requires…
Nicolaï Hoffmann, Aurore Besson, Edouard Cadieu, Matthias Lorthiois + 6 more
With the advent of complete genome assemblies, genome annotation has become essential for the functional interpretation of genomic data. Long-read RNA sequencing (LR-RNAseq) technologies have significantly improved transcriptome annotation by enabling full-length transcript reconstruction for both coding and non-coding…
Alex R. Van Dam, Francisco Hita Garcia
The accelerating biodiversity crisis demands new approaches to taxonomic description that can scale beyond the capacity of professional taxonomists alone. We present the Descriptron-GBIF Annotator, a zero-installation, browser-based tool for morphological annotation of biodiversity specimen images retrieved directly…
Authors not listed
Computational models predicting the sites of metabolism (SOM) of small or- ganic molecules have become invaluable tools for studying and optimizing the metabolic properties of xenobiotics. However, the performance of SOM predic- tors has shown signs of plateauing in recent years, primarily due to the limited…
Chunyan Zhang, Qinglei Zhou, Meng Qiao, Ke Tang + 4 more
'Fudong Liu' 'Qiang Zhang' 'Yifeng Zeng'] Source code summarization (SCS) is a natural language description of source code functionality. It can help developers understand programs and maintain software efficiently. Retrieval-based methods generate SCS by reorganizing terms selected from source code or use SCS of…
Sebastien Lelong, Xinghua Zhou, Cyrus Afrasiabi, Zhongchao Qian + 9 more
To meet the increased need of making biomedical resources more accessible and reusable, Web APIs or web services have become a common way to disseminate knowledge sources. The BioThings APIs are a collection of high-performance, scalable, annotation as a service APIs that automate the integration of biological…
Andrea Ghelfi, Kenta Shirasawa, Sachiko Isobe
The rapid growth of next-generation sequencing (NGS) technology has led to a surge in the determination of whole genome sequences in plants. This has created a need for functional annotation of newly predicted gene sequences in the assembled genomes. To address this, “Hayai-Annotation Plants” was developed as a gene…