19 papers · ranked by Valyu relevance
Qingyuan Liang, Zhao Zhang, Zeyu Sun, Zheng Lin + 8 more
'Yueyi Xiao' 'Yizhou Chen' 'Yuqun Zhang' 'Haotian Zhang' 'Lu Zhang' 'Bin Chen' 'Yingfei Xiong'] Grammar serves as a cornerstone in programming languages and software engineering, providing frameworks to define the syntactic space and program structure. Existing research demonstrates the effectiveness of grammarbased…
Peter Heringer, Daniel Doerr
Pangenome graphs offer a compact and comprehensive representation of genomic diversity, improving tasks such as variant calling, genotyping, and other downstream analyses. Although the underlying graph structures scale sublinearly with the number of haplotypes, the widely used GFA file format suffers from rapidly…
Yihong Dong, Xue Jiang, Yuchen Liu, Ge Li + 1 more
In the process of code generation, it is essential to guarantee the generated code satisfies grammar constraints of programming language (PL). However, neglecting grammar constraints is a fatal drawback of commonly used sequence-based code generation. In this paper, we devise a pushdown automaton (PDA)-based…
Mohammad Jalili Torkamani
Understanding and extracting the grammar of a domainspecific language (DSL) is crucial for various software engineering tasks; however, manually creating these grammars is time-intensive and error-prone. This paper presents Kajal, a novel approach that automatically infers grammar from DSL code snippets by leveraging…
Sanjay Nag, Nabanita Basu, Payal Bose, Samir Kumar Bandyopadhyay + 1 more
'Yunfeng Wu'] Disease prediction using computer-based methods is now an established area of research. The importance of technological intervention is necessary for the better management of disease, as well as to optimize use of limited resources. Various AI-based methods for disease prediction have been documented in…
Cristian Robledo, Francesca Sallicati, Gaël de Chalendar, Marcos Fernández + 4 more
'Marcos Fernández' 'Pablo de Castro' 'Eduardo Martín' 'Javier Gutiérrez' 'Yannis Bouachera'] This paper aims to introduce the innovative work carried out in the Horizon 2020 DECODER project - acronym for “DEveloper COmpanion for Documented and annotatEd code Reference” - (Grant Agreement no. 824231) by linking the…
Richard Apodaca
Despite its widespread use, Simplified Molecular Input Line Entry System (SMILES) remains underspecified. The lack of a detailed specification encourages improvisation by software developers, complicates data standardization efforts, and undermines extension development. Balsa, a reformulation of SMILES, addresses…
Matteo Ciccaglione, Pierciro Caliandro, Alessandro Pellegrini
In this article, we present Tahr, a framework that allows taking attribute grammar specifications and generating a set of software artefacts that can be used programmatically to operate on text compliant with the grammars. Tahr can be used as an algorithmic workbench to test different manipulations of attribute…
Feifei Li, Xiao Chen, Xi Xiao, Xiaoyu Sun + 3 more
'Shaohua Wang' 'Jitao Han'] Black-box context-free grammar inference presents a significant challenge in many practical settings due to limited access to example programs. The state-of-the-art methods, Arvada and Treevada, employ heuristic approaches to generalize grammar rules, initiating from flat parse trees and…
Johannes T. Margraf, Zachary W. Ulissi, Yousung Jung, Karsten Reuter
The discovery of new catalytically active materi- als is one of the holy grails of computational chemistry as it has the potential to accelerate the adoption of renewable energy sources and reduce the energy consumption of chemical industry. Indeed, heterogeneous catalysts are essential for the production of synthetic…
Mohammad Rifat Arefin, Shanto Rahman, Christoph Csallner
Black-box context-free grammar inference is crucial for program analysis, reverse engineering, and security, yet existing tools such as Arvada, TreeVada, and Kedavra struggle with scalability, readability, and accuracy on large, complex languages. We present NatGI, a novel LLM-guided grammar inference framework that…
Muhammad Hammad, Önder Babur, Hamid Abdul Basit, Mark van den Brand + 1 more
'Yilun Shang'] Software developers frequently reuse source code from repositories as it saves development time and effort. Code clones (similar code fragments) accumulated in these repositories represent often repeated functionalities and are candidates for reuse in an exploratory or rapid development. To facilitate…
Morgan Thomas, Mazen Ahmad, Gary Tresadern, Gianni de Fabritiis
SMILES-based generative models are amongst the most robust and successful recent methods used to augment drug design. They are typically used for complete de novo generation, however, scaffold decoration and fragment linking applications are sometimes desirable which requires a different architecture, a different…
Authors not listed
RNA molecules fold into complex three-dimensional structures that determine their function. A wide range of mathematical frameworks, such as chord diagrams, fatgraphs, and context-free grammars, have been used to represent these structures; however, these models have largely been developed from mathematical motivations…
Ying Yin, Yuhai Zhao, Yiming Sun, Chen Chen + 1 more
At present, the explosive growth of software code volume and quantity makes the code review process very labor-intensive and time-consuming. An automated code review model can assist in improving the efficiency of the process. Tufano et al., designed two automated tasks to help improve the efficiency of code review…
Yiwen Zhang, Wei Liu, Fazhong Jiang, Jiquan Ma + 4 more
Large Language Models of the Transformer architecture display great promise in automated code error detection based on their strength in processing sequential data. Nevertheless, their efficacy could be further improved by addressing the inherent weakness in handling structural code dependencies. In response to this…
Melissa Sanabria, Jonas Hirsch, Anna R. Poetsch
Large Language Models (LLMs) on natural language have achieved a level of performance that allows the generation of coherent and syntactically correct text. DNA sequence of genomes follows rules similar to natural language, but a distinguishing factor is the absence of a concept analogous to words. We established…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Pavel Kodytek, Alexandra Bodzas, Jan Zidek, Govind Vashishtha
Continual technological advances associated with the recent automation revolution have tremendously increased the impact of computer technology in the industry. Software development and testing are time-consuming processes, and the current market faces a lack of specialized experts. Introducing automation to this field…