21 papers · ranked by Valyu relevance
Jin Qin, Zihan Liao, Ziyin Zhang, Hang Yu + 2 more
We present C2LLM - Contrastive Code Large Language Models, a family of code embedding models in both 0.5B and 7B sizes. Building upon Qwen-2.5-Coder backbones, C2LLM adopts a Pooling by Multihead Attention (PMA) module for generating sequence embedding from token embeddings, effectively 1) utilizing the LLM's causal…
Benedikt Fein, Gordon Fraser
The trend of embedding source code for machine learning applications also enables new opportunities in learning analytics in programming education, but which code embedding approach is most suitable for learning analytics remains an open question. A common approach to embedding source code lies in treating the code as…
Wangjie Zheng, Yuhan Xie, Jianlei Gu, Hongyu Li + 4 more
Identifying genes associated with rare diseases remains challenging due to the scarcity of patients and the limited statistical power of traditional association methods. Here, we introduce PERADIGM ( Phenotype Embedding similarity-based RAre DIsease Gene Mapping), a novel framework that leverages natural language…
Daria Cherniuk, Nikita Sukhorukov, Gusak, Danil + 5 more
Retrieval-augmented generation has emerged as one of the most effective approaches for code completion, particularly when context from a surrounding repository is essential. However, incorporating context significantly extends sequence length, leading to slower inference—a critical limitation for interactive settings…
Yiwen Zhang, Wei Liu, Fazhong Jiang, Jiquan Ma + 4 more
Large Language Models of the Transformer architecture display great promise in automated code error detection based on their strength in processing sequential data. Nevertheless, their efficacy could be further improved by addressing the inherent weakness in handling structural code dependencies. In response to this…
M. Shahbaz Ismail, Sara Shahzad, Fahmi H. Quradaa, Sajid Anwar
Semantic code clone detection plays an essential role in software maintenance and quality assurance, as it helps uncover fragments of code that express the same logic even when their syntax has been altered or deliberately obfuscated. In this study, we propose a framework that combines hybrid representation learning…
Neusha Javidnia, Ruisi Zhang, Ashish Kundu, Farinaz Koushanfar
—We present SWaRL, a robust and fidelity-preserving watermarking framework designed to protect the intellectual property of code LLM owners by embedding unique and verifiable signatures in the generated output. Existing approaches rely on manually crafted transformation rules to preserve watermarked code functionality…
Istiaq Ahmed Fahad, Mridha Md. Nafis Fuad, Kazi Sakib
Watermarking has become a crucial technique for ensuring provenance and accountability in AI-generated source code. As large language models (LLMs) are increasingly integrated into development workflows, reliable attribution remains challenging. In practice, most developers rely on commercial LLM APIs operating under…
Authors not listed
High-level quantum mechanical (QM) simulations provide accurate electronic information of chemical systems but scale unfavourably with system size, making calculations of applied systems challenging. Hierarchical quantum mechanics in quantum mechanics embedding (QM/QM) addresses this issue by localising the highly…
Niklas Brunn, Sonia Maria Krißmer, Maximilian Frosch, Markus Frick + 2 more
The single-cell literature catalogs cell states as validated marker-gene programs — a sparse, compositional prior. Conventional embedding methods do not leverage this prior and learn cell-state structure de novo from the expression matrix, producing dense dimensions needing post-hoc interpretation and batch correction.…
Yichen Wang, Yijie Lin, Ching-Chun Chang, Chin-Chen Chang + 2 more
With the rapid advancement of the Internet of Medical Things (IoMT), the efficient transmission and management of large-scale medical images in bandwidth- and resource-constrained networks remain critical challenges. This paper proposes a high-payload data hiding method in Absolute Moment Block Truncation Coding…
Melih Peker, Ozcan Ozturk
Selecting a good set of optimization flags requires extensive effort and expert input. While most of the prior research considers using static, spatial, or dynamic features, some of the latest research directly applied deep neural networks to source code. We combined the static features, spatial features, and deep…
Yanshuo Chen, Yuming Zhang, Joshua Li, Boxue Tian + 1 more
Codon optimization involves selecting synonymous codons to match host-specific preferences. It is critical for heterologous expression but remains challenging due to the combinatorial design space. Under long-term evolutionary selection, natural coding sequences are near-optimal compromises between translational…
Beiji Lu
Synonymous codons encode the same amino acid yet are used non-randomly across genomes, a phenomenon with well-documented functional consequences for translation efficiency and mRNA stability. Whether the information embedded in synonymous codon choice is recoverable from the internal representations of in-dependently…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Authors not listed
Recent years have seen a growing interest in machine learning approaches for chemical tasks. The best existing methods focus on building base models that combine molecular graphs (“2D structures”) with atomic coordinates in 3D to predict molecular properties, typically through pre-training followed by fine-tuning on…
André Silva, Han Tu, Martin Monperrus
A coding agent solving a software-engineering task spends dozens of steps reasoning, editing code, and running tests, yet little is known about what the underlying language model internally represents about the program it is working on. We show that the residual streams of language models under coding agents linearly…
Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Zipeng Xie + 5 more
Rendering source code as images offers a promising way to reduce the input costs of Multimodal Large Language Models (MLLMs). Adjusting image resolution can trade visual token cost against content fidelity. However, resolution scaling alone overlooks two sources of inefficiency: blank regions created by line breaks and…
Ramy Khabbaz, Jérémy Mateos, Marc Antonini, Serge Kas Hanna
The biochemical processes underlying DNA data storage, including synthesis, amplification, and sequencing, are inherently noisy. Consequently, base-level insertion, deletion, and substitution (IDS) errors, as well as sequence-level dropouts, occur and pose major challenges for reliable data retrieval. Here we introduce…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…
Authors not listed
Olfaction arises from the interaction of odorants with olfactory receptors, a process shaped by molecular geometry, electron distribution, and conformational preference. We present ConfDENSE, a Set2Set enhanced PointNet model that learns directly from Hirshfeld promolecule electron-density point clouds, preserving full…