20 papers · ranked by Valyu relevance
Jinsheng Ba, Sverrir Thorgeirsson, Zhendong Su
Recent advances in Large Language Models (LLMs) have introduced a new paradigm for software development, where source code is generated directly from natural language prompts. While this paradigm significantly boosts development productivity, building complex, real-world software systems remains challenging because…
Iskander Akhmetov, Timur Saparov, Volkan Duran, Alexander Pak
DNA is often described as the “language of life” because it encodes biological information using nucleotide sequences. Unlike the traditional view focused on codon-to-amino acid mapping in coding regions, the vast non-coding genome reveals complex organizational patterns resembling natural language. This paper outlines…
Hanqi Li, Lu Chen, Kai Yu
As LLMs are increasingly integrated into agentic systems, they must adhere to dynamically defined, machine-interpretable interfaces. We evaluate LLMs as in-context interpreters: given a novel context-free grammar, can LLMs generate syntactically valid, behaviorally functional, and semantically faithful outputs? We…
Feifei Li, Xiao Chen, Xiaoyu Sun, Xi Xiao + 4 more
Grammar inference for complex programming languages remains a significant challenge, as existing approaches fail to scale to realworld datasets within practical time constraints. In our experiments, none of the state-of-the-art tools, including Arvada, Treevada and Kedavra were able to infer grammars for complex…
Yongmin Li, Yihong Dong, Jia Li, Ge Li
LLMs are widely used to generate structured output like source code or JSON. Grammar-constrained decoding (GCD) can guarantee the syntactic validity of the generated output, by masking out tokens that violate rules specified by a context-free grammar. However, the online computational overhead of existing GCD methods…
Matteo Ciccaglione, Pierciro Caliandro, Alessandro Pellegrini
In this article, we present Tahr, a framework that allows taking attribute grammar specifications and generating a set of software artefacts that can be used programmatically to operate on text compliant with the grammars. Tahr can be used as an algorithmic workbench to test different manipulations of attribute…
Weixing Zhang, Bowen Jiang, Rahul Sharma, Regina Hebig + 1 more
In model-driven engineering, metamodel evolution leads to the need to adapt corresponding grammars to maintain consistency, which typically requires tedious manual work. Existing rule-based methods can achieve partial automation but have limitations when handling complex grammar scenarios. This paper proposes a Large…
Lucas Y. Tian, Daniel J. Hanuska, Kedar Garzón Gupta, Yue Liu + 3 more
Humans and other animals can solve new problems, even on the first attempt. This capacity to generate novel problem-solving behavior has been hypothesized to depend on brain mechanisms for recombining units of knowledge using systems of procedural rules, or grammars. Yet, whether and how the brain represents and…
Jieni Hu, David R. Koes, Maria Chikina
Cis-regulatory elements shape gene expression by recruiting transcription factors (TFs) to DNA motifs, yet how motifs cooperate or compete across distances remains poorly understood. Existing deep learning models predict TF binding with high accuracy but rely on computationally intensive post-hoc analyses to infer…
Jesus Antonio Motta, Carolina Fernandez, Maria del Mar Motta
In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide…
Authors not listed
RNA molecules fold into complex three-dimensional structures that determine their function. A wide range of mathematical frameworks, such as chord diagrams, fatgraphs, and context-free grammars, have been used to represent these structures; however, these models have largely been developed from mathematical motivations…
Ghoummaid, Marah, Tchuiev, Vladimir + 6 more
AI-based code generation is increasingly prevalent, with GitHub Copilot estimated to generate 46% of the code on GitHub. Accurately evaluating how well generated code aligns with developer intent remains a critical challenge. Traditional evaluation methods, such as unit tests, are often unscalable and costly. Syntactic…
Zi Wang, Xiaoyu Zhu, Hongqiang Wang, Yichun Yu + 2 more
Formal verification ensures software correctness but faces challenges in kernel specification writing, which is labor-intensive, expertise-dependent, and limited to specific targets. For complex microkernels like seL4, these issues significantly reduce the practicality of formal methods. To address this, we propose…
Yiwen Zhang, Wei Liu, Fazhong Jiang, Jiquan Ma + 4 more
Large Language Models of the Transformer architecture display great promise in automated code error detection based on their strength in processing sequential data. Nevertheless, their efficacy could be further improved by addressing the inherent weakness in handling structural code dependencies. In response to this…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Giuliana Nardacchione, Pierluigi Zoccolotti, Chiara Valeria Marinelli, Nicola Molinaro
Artificial grammar learning (AGL) has frequently been employed to investigate the procedural hypothesis of dyslexia. However, most studies did not distinguish whether performance depended upon the acquisition of grammatical rules (procedural memory), distributional knowledge (statistical learning) or reference to…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
M. Shahbaz Ismail, Sara Shahzad, Fahmi H. Quradaa, Sajid Anwar
Semantic code clone detection plays an essential role in software maintenance and quality assurance, as it helps uncover fragments of code that express the same logic even when their syntax has been altered or deliberately obfuscated. In this study, we propose a framework that combines hybrid representation learning…
Hung Q. Vo, Huy Q. Vo, Son T. Ly, Zhihao Wan + 5 more
Conventional tissue image analysis software provides foundational capabilities for cellular analysis, including segmentation, basic morphological feature extraction, and spatial organization analysis. However, these tools often require manual intervention and are not well integrated with code-driven automation…
Ming-Feng Yeh, Ching-Chuan Luo, Cheng-Lin Lu, Nianbo Liu
Smart manufacturing relies on programmable logic controllers (PLCs) that translate sensor inputs into actuator commands. Generating PLC programs in legacy textual languages such as Mitsubishi FX-series Instruction List (IL) remains an expert-only task, and IL’s deprecation in IEC 61131-3 Edition 3.0 leaves it…