23 papers · ranked by Valyu relevance
Kehao Mao, Baokun Hu, Ruixin Lin, Zewen Li + 2 more
Automated programming has become a powerful tool for solving real-world problems. Code generation, in particular, plays a key role in improving developer productivity and reducing the entry barrier to software development. Recent advances in large language models (LLMs) have significantly improved program synthesis…
Zhao, Qianhui, Zhang, Li + 18 more
In recent years, Large Language Models (LLMs) have achieved remarkable progress in automated code generation. In real-world software engineering, the growing demand for rapid iteration and continuous delivery underscores the importance of project-level code generation, where LLMs are expected to generate complete…
Wang Bin, Li Hui, Liu Aofan, Yang Botao + 7 more
Security in code generation remains a pivotal challenge when applying large language models (LLMs). This paper introduces RefleXGen, an innovative method that significantly enhances code security by integrating Retrieval-Augmented Generation (RAG) techniques with guided selfreflection mechanisms inherent in LLMs.…
Zheng Fang, Yihong Dong, Lili Mou, Dongming Jin + 2 more
Large Language Models (LLMs) have shown strong capabilities in code generation, but their adherence to fine-grained user intent with multiple constraints remains a significant challenge. Our empirical analysis reveals two key observations: 1) Model performance deteriorates quickly as the number of constraints in the…
Sicong Liu, Yanxian Huang, Mingwei Liu, Ting Chen + 5 more
—Code generation tasks aim to automate the conversion of user requirements into executable code, significantly reducing manual development efforts and enhancing software productivity. The emergence of large language models (LLMs) has significantly advanced code generation, though their efficiency is still impacted by…
A. Q. Liu, Haoxuan Li, Bin Wang, Ao Yang + 1 more
—Code generation models based on large language models (LLMs) have gained wide adoption, but challenges remain in ensuring safety, accuracy, and controllability, especially for complex tasks. Existing methods often lack dynamic integration of external tools, transparent reasoning, and user control over safety. To…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Dawei Yuan, Guojun Liang, Tingting Li, Suping Liu
We present a reinforcement learning framework that enhances natural language queries to improve DeepSeek code generation. A parametric refiner (Qwen with LoRA) is trained via REINFORCE while the generator remains fixed, using a scalar reward that can combine text similarity (BLEU-4, ROUGE-L, F1, Overlap) with execution…
Zi Wang, Xiaoyu Zhu, Hongqiang Wang, Yichun Yu + 2 more
Formal verification ensures software correctness but faces challenges in kernel specification writing, which is labor-intensive, expertise-dependent, and limited to specific targets. For complex microkernels like seL4, these issues significantly reduce the practicality of formal methods. To address this, we propose…
Shiqi Kuang, Zhao Tian, Tao Xiao, Dong Wang + 1 more
Large language models (LLMs) have achieved remarkable progress in code generation, largely driven by the availability of high-quality code datasets for effective training. To further improve data quality, numerous training data optimization techniques have been proposed; however, their overall effectiveness has not…
Ruofan Gao, Amjed Tahir, Peng Liang, Teo Sušnjak + 1 more
Developers are widely using AI code-generation models, aiming to increase productivity and efficiency. However, there are also quality concerns regarding the AI-generated code. The generated code is produced by models trained on publicly available code, which are known to contain bugs and quality issues. Those issues…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Niaz Bahar Chowdhury, August George, Sumit Purohit, Angela Cintolesi + 17 more
Genome-scale metabolic models (GEMs) are powerful tools for predicting cellular phenotypes and guiding microbial strain engineering, yet broad adoption remains challenging due to the computational expertise required. To overcome that, we present ChatGEM, an agentic platform that enables interactive GEM simulation…
Ming-Feng Yeh, Ching-Chuan Luo, Cheng-Lin Lu, Nianbo Liu
Smart manufacturing relies on programmable logic controllers (PLCs) that translate sensor inputs into actuator commands. Generating PLC programs in legacy textual languages such as Mitsubishi FX-series Instruction List (IL) remains an expert-only task, and IL’s deprecation in IEC 61131-3 Edition 3.0 leaves it…
Yanshuo Chen, Yuming Zhang, Joshua Li, Boxue Tian + 1 more
Codon optimization involves selecting synonymous codons to match host-specific preferences. It is critical for heterologous expression but remains challenging due to the combinatorial design space. Under long-term evolutionary selection, natural coding sequences are near-optimal compromises between translational…
Hung Q. Vo, Huy Q. Vo, Son T. Ly, Zhihao Wan + 5 more
Conventional tissue image analysis software provides foundational capabilities for cellular analysis, including segmentation, basic morphological feature extraction, and spatial organization analysis. However, these tools often require manual intervention and are not well integrated with code-driven automation…
Xiaojian Liu, Yangyang Zhang, Chee Wei Tan, Wenyi Zhang
Code coverage-guided unit test generation (CGTG) and large language model-based test generation (LLMTG) are two principal approaches for the generation of unit tests. Each of these approaches has its inherent advantages and drawbacks. Tests generated by CGTG have been shown to exhibit high code coverage and high…
Ihsan Tolga Medeni, Metehan Ünal, Roberto Galizi, Bryan Bartley + 4 more
Large language models have transformed software engineering practices. However, generated artefacts are not always developer-friendly and may partially meet complex requirements. As the need to standardise, integrate, and develop tools in engineering biology increases, novel approaches are needed to create and maintain…
Authors not listed
We provide an overview of core molSimplify functionality and recent updates that enhance its capabilities for automated molecular and materials modeling. We describe the mol3D and atom3D classes, which store atomic and bonding information for a wide range of functions, including reading, modifying, and characterizing…
Ksenia Sokolova, Dmitri Kosenkov, Keerthana Nallamotu, Sanketh Vedula + 3 more
The growing availability of biological data resources has transformed research, yet their effective use remains challenging: selecting appropriate sources requires domain knowledge, data are fragmented across databases, and synthesizing results into reliable conclusions is labor-intensive. Although large language…
Authors not listed
TurtleMol is an open-source Python package that aims to help users generate large, complex molec- ular systems. In the current version, users can generate systems by filling volumes defined by basic geometric shapes (e.g. cube, sphere), or by shapes of arbitrary gemoetries defined meshes created in other software (such…
Authors not listed
Here, we present MolPic, an open-source Python-based software that can be used to generate high-resolution, publication-quality molecular figures directly from compound names or SMILES strings. MolPic supports single-molecule rendering, batch processing, and automated multi-panel 2D figure generation, which are…
Authors not listed
Incorporating prior domain knowledge into Bayesian optimization (BO) remains difficult for statistical methods, which also typically suffer from limited interpretability. Large language models (LLMs) offer complementary strengths in reasoning and knowledge integration, but it remains unclear when and how they improve…