20 papers · ranked by Valyu relevance
Qingyuan Liang, Zhao Zhang, Zeyu Sun, Zheng Lin + 8 more
'Yueyi Xiao' 'Yizhou Chen' 'Yuqun Zhang' 'Haotian Zhang' 'Lu Zhang' 'Bin Chen' 'Yingfei Xiong'] Grammar serves as a cornerstone in programming languages and software engineering, providing frameworks to define the syntactic space and program structure. Existing research demonstrates the effectiveness of grammarbased…
Peter Heringer, Daniel Doerr
Pangenome graphs offer a compact and comprehensive representation of genomic diversity, improving tasks such as variant calling, genotyping, and other downstream analyses. Although the underlying graph structures scale sublinearly with the number of haplotypes, the widely used GFA file format suffers from rapidly…
Yihong Dong, Xue Jiang, Yuchen Liu, Ge Li + 1 more
In the process of code generation, it is essential to guarantee the generated code satisfies grammar constraints of programming language (PL). However, neglecting grammar constraints is a fatal drawback of commonly used sequence-based code generation. In this paper, we devise a pushdown automaton (PDA)-based…
Cristian Robledo, Francesca Sallicati, Gaël de Chalendar, Marcos Fernández + 4 more
'Marcos Fernández' 'Pablo de Castro' 'Eduardo Martín' 'Javier Gutiérrez' 'Yannis Bouachera'] This paper aims to introduce the innovative work carried out in the Horizon 2020 DECODER project - acronym for “DEveloper COmpanion for Documented and annotatEd code Reference” - (Grant Agreement no. 824231) by linking the…
Alvaro Veizaga, Mauricio Alferez, Damiano Torre, Mehrdad Sabetzadeh + 1 more
'Lionel Briand'] Natural language (NL) is pervasive in software requirements specifications (SRSs). However, despite its popularity and widespread use, NL is highly prone to quality issues such as vagueness, ambiguity, and incompleteness. Controlled natural languages (CNLs) have been proposed as a way to prevent…
Oscar Westesson, Ian Holmes, Christos A. Ouzounis
Modeling sequence evolution on phylogenetic trees is a useful technique in computational biology. Especially powerful are models which take account of the heterogeneous nature of sequence evolution according to the “grammar” of the encoded gene features. However, beyond a modest level of model complexity, manual coding…
Richard Apodaca
Despite its widespread use, Simplified Molecular Input Line Entry System (SMILES) remains underspecified. The lack of a detailed specification encourages improvisation by software developers, complicates data standardization efforts, and undermines extension development. Balsa, a reformulation of SMILES, addresses…
Matteo Ciccaglione, Pierciro Caliandro, Alessandro Pellegrini
In this article, we present Tahr, a framework that allows taking attribute grammar specifications and generating a set of software artefacts that can be used programmatically to operate on text compliant with the grammars. Tahr can be used as an algorithmic workbench to test different manipulations of attribute…
Feifei Li, Xiao Chen, Xi Xiao, Xiaoyu Sun + 3 more
'Shaohua Wang' 'Jitao Han'] Black-box context-free grammar inference presents a significant challenge in many practical settings due to limited access to example programs. The state-of-the-art methods, Arvada and Treevada, employ heuristic approaches to generalize grammar rules, initiating from flat parse trees and…
Vadim Zaytsev
| 1 | Introduction | | 1 | | --- | --- | --- | --- | | 2 | Preliminaries | | 1 | | | 2.1 | Background notions | 1 | | | 2.2 | Major contributions in a nutshell | 2 | | | 2.3 | Selected minor contributions | 3 | | | 2.4 | Motivation for this report | 4 | | 3 | | Topics overview | 5 | | | 3.1 | Guided grammar convergence…
Johannes T. Margraf, Zachary W. Ulissi, Yousung Jung, Karsten Reuter
The discovery of new catalytically active materi- als is one of the holy grails of computational chemistry as it has the potential to accelerate the adoption of renewable energy sources and reduce the energy consumption of chemical industry. Indeed, heterogeneous catalysts are essential for the production of synthetic…
Mohammad Rifat Arefin, Shanto Rahman, Christoph Csallner
Black-box context-free grammar inference is crucial for program analysis, reverse engineering, and security, yet existing tools such as Arvada, TreeVada, and Kedavra struggle with scalability, readability, and accuracy on large, complex languages. We present NatGI, a novel LLM-guided grammar inference framework that…
Morgan Thomas, Mazen Ahmad, Gary Tresadern, Gianni de Fabritiis
SMILES-based generative models are amongst the most robust and successful recent methods used to augment drug design. They are typically used for complete de novo generation, however, scaffold decoration and fragment linking applications are sometimes desirable which requires a different architecture, a different…
Authors not listed
RNA molecules fold into complex three-dimensional structures that determine their function. A wide range of mathematical frameworks, such as chord diagrams, fatgraphs, and context-free grammars, have been used to represent these structures; however, these models have largely been developed from mathematical motivations…
Chen Yang, Yan Liu, Changqing Yin, Karsten Keller
Source Code Generation (SCG) is a prevalent research field in the automation software engineering sector that maps specific descriptions to various sorts of executable code. Along with the numerous intensive studies, diverse SCG types that integrate different scenarios and contexts continue to emerge. As the ultimate…
R. Daniel Kortschak, David L. Adelson
bíogo is a framework designed to ease development and maintenance of computationally intensive bioinformatics applications. The library is written in the Go programming language, a garbage-collected, strictly typed compiled language with built in support for concurrent processing, and performance comparable to C and…
Melissa Sanabria, Jonas Hirsch, Anna R. Poetsch
Large Language Models (LLMs) on natural language have achieved a level of performance that allows the generation of coherent and syntactically correct text. DNA sequence of genomes follows rules similar to natural language, but a distinguishing factor is the absence of a concept analogous to words. We established…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Pavel Kodytek, Alexandra Bodzas, Jan Zidek, Govind Vashishtha
Continual technological advances associated with the recent automation revolution have tremendously increased the impact of computer technology in the industry. Software development and testing are time-consuming processes, and the current market faces a lack of specialized experts. Introducing automation to this field…
Jacqueline A Jansen, Artür Manukyan, Nour Al Khoury, Altuna Akalin
Data analysis is constrained by a shortage of skilled experts, particularly in biology, where detailed data interpretation is vital for understanding complex biological processes and developing new treatments and diagnostics. To address this, we developed mergen, an R package that leverages Large Language Models (LLMs)…