22 papers · ranked by Valyu relevance
Jin Qin, Zihan Liao, Ziyin Zhang, Hang Yu + 2 more
We present C2LLM - Contrastive Code Large Language Models, a family of code embedding models in both 0.5B and 7B sizes. Building upon Qwen-2.5-Coder backbones, C2LLM adopts a Pooling by Multihead Attention (PMA) module for generating sequence embedding from token embeddings, effectively 1) utilizing the LLM's causal…
Yang Yang, Li Kuang, Jiakun Liu, Zhongxin Liu + 2 more
Effective code retrieval is indispensable and it has become an important paradigm to search code in hybrid mode using both natural language and code snippets. Nevertheless, it remains unclear whether existing approaches can effectively leverage such hybrid queries, particularly in cross-language contexts. We conduct a…
Junsong Pu, Yichen Li, Zhuangbin Chen
Graph-based code indexing can improve context retrieval for LLM-based code agents by preserving call chains and dependency relationships that keyword search and similarity retrieval often miss. ABCoder is an open-source framework that parses codebases into a function-level code index called UniAST, but its existing…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Baoyi Wang, Xingliang Wang, Guochang Li, Chen Zhi + 5 more
Repository-level code completion remains challenging for large language models (LLMs), as it requires reasoning over cross-file dependencies while under limited context windows. To address this challenge, prior work has adopted Retrieval-Augmented Generation (RAG) frameworks based on semantic indexing or structureaware…
Leonardo Venuta, Francesco Tosoni, Paolo Ferragina
Semantic code search and clone detection are essential for software development, maintenance, and reuse. This paper evaluates the effectiveness, efficiency, and scalability of contemporary deep learning models for first-stage recall in large-scale code-to-code search engines. Benchmarking across multiple programming…
Ye Fan, Jidong Ge, Chuanyi Li, Liguo Huang + 1 more
While pre-trained models have achieved remarkable success in code search, their multilingual capabilities remain a major hurdle, plagued by data imbalance, cross-lingual semantic interference, and the loss of critical information from existing unified representations like Abstract Syntax Trees (ASTs) or Intermediate…
Shahd Seddik, Fahd Seddik, Iman Saberi, Fatemeh H. Fard + 2 more
Large Language Models (LLMs) excel at code generation but struggle with complex problems. Retrieval-Augmented Generation (RAG) mitigates this issue by integrating external knowledge, yet retrieval models often miss relevant context, and generation models hallucinate with irrelevant data. We propose Programming…
Jin Lu, Ji Li, Zheng Zhang
This paper presents a novel approach to enhancing educational question-answering (Q&A) systems by combining Retrieval-Augmented Generation (RAG) with Large Language Model (LLM) Code Interpreters. Traditional educational Q&A systems face challenges in areas such as knowledge updates, reasoning accuracy, and the handling…
Dawei Yuan, Guojun Liang, Tingting Li, Suping Liu
We present a reinforcement learning framework that enhances natural language queries to improve DeepSeek code generation. A parametric refiner (Qwen with LoRA) is trained via REINFORCE while the generator remains fixed, using a scalar reward that can combine text similarity (BLEU-4, ROUGE-L, F1, Overlap) with execution…
Ming-Feng Yeh, Ching-Chuan Luo, Cheng-Lin Lu, Nianbo Liu
Smart manufacturing relies on programmable logic controllers (PLCs) that translate sensor inputs into actuator commands. Generating PLC programs in legacy textual languages such as Mitsubishi FX-series Instruction List (IL) remains an expert-only task, and IL’s deprecation in IEC 61131-3 Edition 3.0 leaves it…
Soham Ghosh, Gaurav Mittal
Large language models (LLMs) are powerful in language understanding and content generation but frequently fall short of technical accuracy when they are applied to engineering code, standards, and design documents. To mitigate this, we are seeing the emergence of Retrieval-Augmented Generation (RAG) models that ground…
Niaz Bahar Chowdhury, August George, Sumit Purohit, Angela Cintolesi + 17 more
Genome-scale metabolic models (GEMs) are powerful tools for predicting cellular phenotypes and guiding microbial strain engineering, yet broad adoption remains challenging due to the computational expertise required. To overcome that, we present ChatGEM, an agentic platform that enables interactive GEM simulation…
Zi Wang, Xiaoyu Zhu, Hongqiang Wang, Yichun Yu + 2 more
Formal verification ensures software correctness but faces challenges in kernel specification writing, which is labor-intensive, expertise-dependent, and limited to specific targets. For complex microkernels like seL4, these issues significantly reduce the practicality of formal methods. To address this, we propose…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Longsheng Jiang, Yanlin Zhu, Jia Liu
Working memory (WM) stores information after sensory input disappears and later retrieves it in a task-relevant format, but the mechanism unifying storage and retrieval remains unclear. Here we combine neural geometry analyses of macaque dorsolateral prefrontal cortex activity during a visuospatial…
Wang Lingling, Sun Shijie, Liu Xueyi
Generative artificial intelligence is now embedded in everyday university work, yet most evidence on its educational value still comes from surveys of student perception rather than from what students do. We reanalyzed the public log of Prof. Leodar, a course-specific retrieval-augmented chatbot deployed in an…
María Peña Fernández, Lara Lloret Iglesias, Jesús Marco de Lucas
How much machinery does a network need to memorize and recall discrete sequences when constrained to a biologically plausible substrate? We address this question using 50 short monophonic melodies in 4/4, used only as a controlled sequence-memory benchmark. Each beat is encoded with two clean one-hot populations – a…
Matthew B. Bair, Nicole M. Long
It is critical to identify which factors induce specific brain states as these large-scale patterns of coordinated neural activity drive downstream processing and behavior. The retrieval state, a brain state engaged when attempting to retrieve the past, is thought to specifically support episodic memory, remembering…
Ibrahim Nawaz, Parv Agarwal, Thomas Heinis
DNA storage is a developing field that uses DNA to archive digital data owing to its superior information density and stability. Although DNA storage has been performed on a significant scale, challenges arise from the synthesis and sequencing of data-encoded oligonucleotides. Synthesis of DNA introduces significant…
Andrew Coristine, Ebenezer Aquisman Asare, Sherif Elmitwalli, Chunwei Ma + 1 more
Background Clinical retrieval-augmented generation depends on embedding models. A companion study found that non-retrieval-trained encoders underperformed retrieval-trained general-purpose embeddings and produced near-degenerate embedding geometry, but did not localize the architectural origin, separate training-domain…
Authors not listed
Agentic artificial intelligence (AI) is poised to redefine how science is conducted, automating not just data analysis but the entire research lifecycle, from hypothesis generation to validation. Yet most current AI agents remain domain-bound, tailored to specific applications such as materials synthesis or quantum…