12 papers · ranked by Valyu relevance
Nikita Pavlichenko, Iurii Nazarov, Ivan Dolgov, Ekaterina Garanina + 9 more
We present the Mellum models family, open-weight code completion models designed for interactive use in JetBrains IDEs. Mellums have 4B parameters, adopt a Llama-style architecture, and are pretrained on 4T tokens of permissively licensed, multi-language code. Our studies show that (i) careful data curation and staged…
Daria Cherniuk, Nikita Sukhorukov, Gusak, Danil + 5 more
Retrieval-augmented generation has emerged as one of the most effective approaches for code completion, particularly when context from a surrounding repository is essential. However, incorporating context significantly extends sequence length, leading to slower inference—a critical limitation for interactive settings…
Jiajun Zhang, Zeyu Cui, Jiaxi Yang, Lei Zhang + 8 more
The dominant Fill-in-the-Middle (FIM) paradigm for code completion is constrained by its rigid inability to correct contextual errors and reliance on unaligned, insecure Base models. While Chat LLMs offer safety and Agentic workflows provide flexibility, they suffer from performance degradation and prohibitive latency…
Omar Abedelkader, Stéphane Ducasse, Oleksandr Zaitsev, Romain Robbes + 1 more
Pharo offers a sophisticated completion engine based on semantic heuristics, which coordinates specific fetchers within a lazy architecture. These heuristics can be recomposed to support various activities (e.g., live programming or history usage navigation). While this system is powerful, it does not account for the…
Rahman, Imranur, Md Rayhanur Rahman
—Code completion can help developers improve efficiency and ease the development lifecycle. Although code completion is available in modern integrated development environments (IDEs), research lacks in determining what makes a good context for code completion based on the information available to the IDEs for the large…
Xinkui Zhao, R. Liu, Yifan Zhang, Zhi Chen + 6 more
As code completion task from function-level to repository-level, leveraging contextual information from large-scale codebases becomes a core challenge. However, existing retrieval-augmented generation (RAG) methods typically treat code as plain natural language, relying primarily on shallow semantic matching while…
Kilian Kier, Alessandro Giagnorio, Omar AbedelKader, Oleksandr Zaitsev + 4 more
Large Language Models (LLMs) unlocked new possibilities in automated code writing, becoming the backbone of most code completion tools. While LLMs excel in mainstream languages, they often lack support for the so-called low-resource languages where training data is scarce. As a result, these languages lag behind in the…
Liang Zhu, Haolin Chen, Lidong Zhao, Xian Wu
While Large Language Models (LLMs) have demonstrated exceptional proficiency in code completion, they typically adhere to a Hard Completion (HC) paradigm, compelling the generation of fully concrete code even amidst insufficient context. Our analysis of 3 million real-world interactions exposes the limitations of this…
Kishanthan Thangarajah, Boyuan Chen, Ahmed E. Hassan
Enterprises want AI code completion that is both high-quality and private, but they face a tension: proprietary models yield better results yet risk exposing proprietary code, while self-hosting large models is expensive and hard to maintain. As a lighter alternative, small CodeLLMs (1B-3B) can run on a developer's…
Mehdi Elkolei, Omar Abedelkader, Stéphane Ducasse
Complishon is Pharo's context-aware code completion engine, built on AST analysis, lazy candidate generation, and filter-based candidate selection. Its existing design already provides strong semantic completion, but several practical limits remain. Strict prefix matching is sensitive to small typing errors, framework…
Jenny T. Liang, Mihika Bairathi, Wayne Chi, Ameet Talwalkar + 2 more
Imperfections in AI-generated code require that software developers modify the generated code manually, or by re-prompting an AI programming assistant. Manual code edits provide more realistic and granular information on editing behavior than Git commits, which only contain final successful code snippets. Yet, due to a…
Benedikt Fein, Gordon Fraser
The trend of embedding source code for machine learning applications also enables new opportunities in learning analytics in programming education, but which code embedding approach is most suitable for learning analytics remains an open question. A common approach to embedding source code lies in treating the code as…