21 papers · ranked by Valyu relevance
Yihong Dong, Xue Jiang, Jiaru Qian, Wang Tian + 2 more
—Code generation agents powered by large language models (LLMs) are revolutionizing the software development paradigm. Distinct from previous code generation techniques, code generation agents are characterized by three core features. 1) Autonomy: the ability to independently manage the entire workflow, from task…
Man-Fai Wong, Shangxin Guo, Ching-Nam Hang, Siu-Wai Ho + 2 more
'Chee-Wei Tan' 'Lei Wang'] This paper provides a comprehensive review of the literature concerning the utilization of Natural Language Processing (NLP) techniques, with a particular focus on transformer-based large language models (LLMs) trained using Big Code, within the domain of AI-assisted programming tasks. LLMs…
Huy Le, Phong Nguyen, Hao Do, Tuan Nguyen + 3 more
Automatic code generation has long been regarded as a critical and aspirational goal in Software Engineering research [1, 2, [3]]. Its core objective is to translate userprovided requirements—often expressed in natural language—into executable source code, thereby streamlining the software development process. This…
Naizhu Jin, Zhong Li, Tian Zhang, Qingkai Zeng
—With the rapid development of code intelligence, the application of multiple programming languages is becoming increasingly widespread. However, most existing code generation models mainly focus on a single or a few programming languages, resulting in unsatisfactory performance in a multilingual environment.…
Ekaterina Trofimova, Emil Sataev, Andrey Ustyuzhanin, Xiangjie Kong
In the ever-evolving landscape of machine learning, seamless translation of natural language descriptions into executable code remains a formidable challenge. This article introduces Linguacodus, an innovative framework designed to tackle this challenge by deploying a dynamic pipeline that iteratively transforms…
Pengcheng Yin, Graham Neubig
We consider the problem of parsing natural language descriptions into source code written in a general-purpose programming language like Python. Existing datadriven methods treat this problem as a language generation task without considering the underlying syntax of the target programming language. Informed by previous…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Xinyi He, Jiaru Zou, Yun Lin, Mengyu Zhou + 3 more
and Correctness Testing Authors: ['Xinyi He' 'Jiaru Zou' 'Yun Lin' 'Mengyu Zhou' 'Han Shi' 'Zejian Yuan' 'Dongmei Zhang'] Large Language Models (LLMs) have revolutionized code generation ability by converting natural language descriptions into executable code. However, generating complex code within real-world…
Chen Yang, Yan Liu, Changqing Yin, Karsten Keller
Source Code Generation (SCG) is a prevalent research field in the automation software engineering sector that maps specific descriptions to various sorts of executable code. Along with the numerous intensive studies, diverse SCG types that integrate different scenarios and contexts continue to emerge. As the ultimate…
Pavel Kodytek, Alexandra Bodzas, Jan Zidek, Govind Vashishtha
Continual technological advances associated with the recent automation revolution have tremendously increased the impact of computer technology in the industry. Software development and testing are time-consuming processes, and the current market faces a lack of specialized experts. Introducing automation to this field…
Yabing Zhu, Yanfeng Zhang, Huili Yang, Fangjing Wang
We propose GANCoder, an automatic programming approach based on Generative Adversarial Networks (GAN), which can generate the same functional and logical programming language codes conditioned on the given natural language utterances. The adversarial training between generator and discriminator helps generator learn…
Zi Wang, Xiaoyu Zhu, Hongqiang Wang, Yichun Yu + 2 more
Formal verification ensures software correctness but faces challenges in kernel specification writing, which is labor-intensive, expertise-dependent, and limited to specific targets. For complex microkernels like seL4, these issues significantly reduce the practicality of formal methods. To address this, we propose…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Jacqueline A Jansen, Artür Manukyan, Nour Al Khoury, Altuna Akalin
Data analysis is constrained by a shortage of skilled experts, particularly in biology, where detailed data interpretation is vital for understanding complex biological processes and developing new treatments and diagnostics. To address this, we developed mergen, an R package that leverages Large Language Models (LLMs)…
Huifang Ma, Zhicheng Ji
Large language models have shown remarkable capabilities in algorithm design, but their effectiveness in solving data science challenges remains poorly understood. We conducted a classroom experiment in which graduate students used large language models (LLMs) to solve biomedical data science challenges on Kaggle.…
Authors not listed
Accelerating computational materials science relies not only on hardware advances but also on software that increases the ease of working with the relevant abstractions. Creation and manipulation of crystal structures is a part of many routine materials science workflows. In this work, we demonstrate how fine tuning…
Balaji Kumar, Supreet Saini
Many theories have been proposed attempting to explain the origin of the genetic code. While strong reasons remain to believe that the genetic code evolved as a frozen accident, at least for the first few amino acids, other theories remain viable. In this work, we test the optimality of the standard genetic code…
Ihsan Tolga Medeni, Metehan Ünal, Roberto Galizi, Bryan Bartley + 4 more
Large language models have transformed software engineering practices. However, generated artefacts are not always developer-friendly and may partially meet complex requirements. As the need to standardise, integrate, and develop tools in engineering biology increases, novel approaches are needed to create and maintain…
W. B. Langdon
Grow and graft genetic programming (GGGP) can automatically evolve an existing state-of-the art program to give more accurate predictions of the secondary structures adapted by RNA molecules using their base sequence alone. That is, genetic improvement (GI) can make functional as well as non-functional source code…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Antony M Jose
Our tremendous progress in understanding living things can mask our ignorance of the information needed to perpetuate life. As the basic unit of life, cells hold information in two distinct forms: in the stable DNA sequence that is faithfully replicated during cell divisions, and in the changing arrangement of…