19 papers · ranked by Valyu relevance
Nikita Pavlichenko, Iurii Nazarov, Ivan Dolgov, Ekaterina Garanina + 9 more
We present the Mellum models family, open-weight code completion models designed for interactive use in JetBrains IDEs. Mellums have 4B parameters, adopt a Llama-style architecture, and are pretrained on 4T tokens of permissively licensed, multi-language code. Our studies show that (i) careful data curation and staged…
Omar Abedelkader, Stéphane Ducasse, Oleksandr Zaitsev, Romain Robbes + 1 more
Pharo offers a sophisticated completion engine based on semantic heuristics, which coordinates specific fetchers within a lazy architecture. These heuristics can be recomposed to support various activities (e.g., live programming or history usage navigation). While this system is powerful, it does not account for the…
Daria Cherniuk, Nikita Sukhorukov, Gusak, Danil + 5 more
Retrieval-augmented generation has emerged as one of the most effective approaches for code completion, particularly when context from a surrounding repository is essential. However, incorporating context significantly extends sequence length, leading to slower inference—a critical limitation for interactive settings…
Rahman, Imranur, Md Rayhanur Rahman
—Code completion can help developers improve efficiency and ease the development lifecycle. Although code completion is available in modern integrated development environments (IDEs), research lacks in determining what makes a good context for code completion based on the information available to the IDEs for the large…
Mehdi Elkolei, Omar Abedelkader, Stéphane Ducasse
Complishon is Pharo's context-aware code completion engine, built on AST analysis, lazy candidate generation, and filter-based candidate selection. Its existing design already provides strong semantic completion, but several practical limits remain. Strict prefix matching is sensitive to small typing errors, framework…
Kilian Kier, Alessandro Giagnorio, Omar AbedelKader, Oleksandr Zaitsev + 4 more
Large Language Models (LLMs) unlocked new possibilities in automated code writing, becoming the backbone of most code completion tools. While LLMs excel in mainstream languages, they often lack support for the so-called low-resource languages where training data is scarce. As a result, these languages lag behind in the…
Kehao Mao, Baokun Hu, Ruixin Lin, Zewen Li + 2 more
Automated programming has become a powerful tool for solving real-world problems. Code generation, in particular, plays a key role in improving developer productivity and reducing the entry barrier to software development. Recent advances in large language models (LLMs) have significantly improved program synthesis…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Tadanobu Chuyo Kamijo, Naoki Nakajima, Takeshi Aihara
The dentate gyrus (DG) decorrelates entorhinal inputs (pattern separation); area CA3 completes partial cues via recurrent autoassociation. The density of CA3 recurrent connectivity is contested, with estimates from ∼0.9% (10) to ∼9–11% (30). We ask how completion depends on recurrent connectivity (C_RC_) and whether…
Yiwen Zhang, Wei Liu, Fazhong Jiang, Jiquan Ma + 4 more
Large Language Models of the Transformer architecture display great promise in automated code error detection based on their strength in processing sequential data. Nevertheless, their efficacy could be further improved by addressing the inherent weakness in handling structural code dependencies. In response to this…
Authors not listed
A framework for catalysis based on categorical aperture selection rather than temporal acceleration is presented. Traditional catalysis theory describes catalysts as agents that accelerate reactions by lowering activation energies, implicitly treating time as the fundamental variable and reaction rate enhancement as…
Hong-Jie Dai, Zheng-Hao Li, An-Tai Lu, Min-I Su + 7 more
Reliable ICD-10-CM coding remains a major operational burden in hospitals, and the real-world performance of AI systems for this task is poorly understood. We developed and deployed a modular, clinically grounded pipeline that combines principled base-model selection, redundancy-aware training, and HL7-aligned section…
Tom Pollard, Thomas Sounack, Catherine A. Gao, Leo Anthony Celi + 5 more
Introduction The Transparent Reporting of a multivariable prediction model of Individual Prognosis Or Diagnosis (TRIPOD) statement was published to improve the reporting and critical appraisal of prediction model studies for diagnosis and prognosis. This paper describes the processes and methods that will be used to…
Huanqiu Zhang, Israel Nelken, Tatyana Sharpee
Deciphering the neural code requires identifying its fundamental symbols or code-words. Neural activity is usually interpreted either as a rate code – based on average spike counts – or as a temporal code, which distinguishes patterns with identical counts. Yet, the symbols of the code remain undefined. Here we show…
Jonathan Adams, Luca Serrière, Maximilian Jonathan Kothen, Maria Barbara Smorczewska + 2 more
In amodal completion observers perceive complete objects despite partial occlusion. When two object parts are divided by an occluder, completion can result in perceiving one or two objects. This phenomenon involves both lower-level cues (e.g., symmetry, contour continuity) and higher-level cues (e.g., prior knowledge).…
Jeremy Li, Alex Rubinsteyn, Sergey Feldman, Timothy O’Donnell + 18 more
Scientific computing has become a central component of modern scientific discovery. Yet many computational tools are developed by small, specialized teams under incentives that encourage the release of rapidly prototyped tooling without commensurate attention to engineering concerns, including performance and…
M. Shahbaz Ismail, Sara Shahzad, Fahmi H. Quradaa, Sajid Anwar
Semantic code clone detection plays an essential role in software maintenance and quality assurance, as it helps uncover fragments of code that express the same logic even when their syntax has been altered or deliberately obfuscated. In this study, we propose a framework that combines hybrid representation learning…
Ksenia Sokolova, Dmitri Kosenkov, Keerthana Nallamotu, Sanketh Vedula + 3 more
The growing availability of biological data resources has transformed research, yet their effective use remains challenging: selecting appropriate sources requires domain knowledge, data are fragmented across databases, and synthesizing results into reliable conclusions is labor-intensive. Although large language…