16 papers · ranked by Valyu relevance
Jin Qin, Zihan Liao, Ziyin Zhang, Hang Yu + 2 more
We present C2LLM - Contrastive Code Large Language Models, a family of code embedding models in both 0.5B and 7B sizes. Building upon Qwen-2.5-Coder backbones, C2LLM adopts a Pooling by Multihead Attention (PMA) module for generating sequence embedding from token embeddings, effectively 1) utilizing the LLM's causal…
Benedikt Fein, Gordon Fraser
The trend of embedding source code for machine learning applications also enables new opportunities in learning analytics in programming education, but which code embedding approach is most suitable for learning analytics remains an open question. A common approach to embedding source code lies in treating the code as…
Rui Xu, Jiawei Chen, Zhaoxia Yin, Cong Kong + 2 more
The widespread use of large language models (LLMs) and open-source code has raised ethical and security concerns regarding the distribution and attribution of source code, including unauthorized redistribution, license violations, and misuse of code for malicious purposes. Watermarking has emerged as a promising…
Daria Cherniuk, Nikita Sukhorukov, Gusak, Danil + 5 more
Retrieval-augmented generation has emerged as one of the most effective approaches for code completion, particularly when context from a surrounding repository is essential. However, incorporating context significantly extends sequence length, leading to slower inference—a critical limitation for interactive settings…
Liliana Hotsko, Yinxi Li, Yuntian Deng, Pengyu Nie
Code language models need repository-level context to resolve imports, APIs, and project conventions. Existing methods inject this knowledge as long inputs (retrieved through RAG or dependency analysis) or through per-repository fine-tuning and LoRA -- costly at repository scale and brittle to evolving codebases. We…
Neusha Javidnia, Ruisi Zhang, Ashish Kundu, Farinaz Koushanfar
—We present SWaRL, a robust and fidelity-preserving watermarking framework designed to protect the intellectual property of code LLM owners by embedding unique and verifiable signatures in the generated output. Existing approaches rely on manually crafted transformation rules to preserve watermarked code functionality…
Jiahui Geng, Qing Li, Fengyu Cai, Fakhri Karray
Code search, framed as information retrieval (IR), underpins modern software engineering and increasingly powers retrieval-augmented generation (RAG), improving code discovery, reuse, and the reliability of LLM-based coding. Yet existing code IR models remain largely text-centric and often overlook the visual and…
Istiaq Ahmed Fahad, Mridha Md. Nafis Fuad, Kazi Sakib
Watermarking has become a crucial technique for ensuring provenance and accountability in AI-generated source code. As large language models (LLMs) are increasingly integrated into development workflows, reliable attribution remains challenging. In practice, most developers rely on commercial LLM APIs operating under…
Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Zipeng Xie + 5 more
Rendering source code as images offers a promising way to reduce the input costs of Multimodal Large Language Models (MLLMs). Adjusting image resolution can trade visual token cost against content fidelity. However, resolution scaling alone overlooks two sources of inefficiency: blank regions created by line breaks and…
Fuwei Zhang, Yanzhao Zhang, Mingxin Li, Dingkun Long + 4 more
Code retrieval is becoming central to coding agents, but agentic coding requires more than matching a natural-language query to an isolated snippet. Given a user request, a coding agent needs to navigate a concrete repository state, locate relevant files and functions, gather supporting context, and filter similar…
Po-Han Cheng, Chia-Mu Yu, Ying-Dar Lin, Yu-Sung Wu + 1 more
Code large language models increasingly retrieve external code context from repositories, documentation, issue threads, and coding-agent environments, creating an indirect prompt-injection surface where attackers hide instructions in comments, strings, identifiers, or decoy code. We propose CodeSentinel, a three-layer…
Myeongsoo Kim, Joe Hsu, Dingmin Wang, Shweta Garg + 2 more
LLM-based code agents treat repositories as unstructured text, applying edits through brittle string matching that frequently fails due to formatting drift or ambiguous patterns. We propose reframing the codebase as a structured action space where agents operate on named AST entities rather than text spans. Our…
André Silva, Han Tu, Martin Monperrus
A coding agent solving a software-engineering task spends dozens of steps reasoning, editing code, and running tests, yet little is known about what the underlying language model internally represents about the program it is working on. We show that the residual streams of language models under coding agents linearly…
Jiajun Zhang, Zeyu Cui, Jiaxi Yang, Lei Zhang + 8 more
The dominant Fill-in-the-Middle (FIM) paradigm for code completion is constrained by its rigid inability to correct contextual errors and reliance on unaligned, insecure Base models. While Chat LLMs offer safety and Agentic workflows provide flexibility, they suffer from performance degradation and prohibitive latency…
Ankit Gupta, Aditya Prasad, Rameswar Panda
Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats. While synthetic data has proven transformative for language models, code remains largely unexplored beyond limited quality improvements. We present CodeAlchemy, a synthetic data generation framework that transforms…
Kevin Pulo
Source code is almost universally edited as plain text. However, the mismatch between the syntactic and semantic requirements of valid and correct code, and the unconstrained text editing process trying to produce it, introduces friction that degrades the programming task. It is also increasingly costly in the era of…