17 papers · ranked by Valyu relevance
Benedikt Fein, Gordon Fraser
The trend of embedding source code for machine learning applications also enables new opportunities in learning analytics in programming education, but which code embedding approach is most suitable for learning analytics remains an open question. A common approach to embedding source code lies in treating the code as…
Jin Qin, Zihan Liao, Ziyin Zhang, Hang Yu + 2 more
We present C2LLM - Contrastive Code Large Language Models, a family of code embedding models in both 0.5B and 7B sizes. Building upon Qwen-2.5-Coder backbones, C2LLM adopts a Pooling by Multihead Attention (PMA) module for generating sequence embedding from token embeddings, effectively 1) utilizing the LLM's causal…
M. Shahbaz Ismail, Sara Shahzad, Fahmi H. Quradaa, Sajid Anwar
Semantic code clone detection plays an essential role in software maintenance and quality assurance, as it helps uncover fragments of code that express the same logic even when their syntax has been altered or deliberately obfuscated. In this study, we propose a framework that combines hybrid representation learning…
Pablo Millan Arias, Niousha Sadjadi, Monireh Safari, ZeMing Gong + 9 more
Inspired by Bidirectional Encoder Representations from Transformers (BERT)-like models, which convert sequence inputs into meaningful embedding vectors, BarcodeBERT is designed to encode DNA barcodes into informative embedding vectors for fast and effective comparisons. This architecture’s main building block is the…
Syed Mehedi Hasan Nirob, Shamim Ehsan, Moqsadur Rahman, Summit Haque
—Large language models (LLMs) have made it remarkably easy to synthesize plausible source code from natural language prompts. While this accelerates software development and supports learning, it also raises new risks for academic integrity, authorship attribution, and responsible AI use. This paper investigates the…
Weiye Li, Wenyi Tang
Source Code Model (SCM) aims to learn the proper embeddings from source codes, demonstrating significant success in various software engineering or security tasks. The recent explosive development of Large Language Model (LLM) extends the family of SCMs, bringing LLMs for code (LLM4Code) that revolutionize development…
Jiahui Geng, Qing Li, Fengyu Cai, Fakhri Karray
Code search, framed as information retrieval (IR), underpins modern software engineering and increasingly powers retrieval-augmented generation (RAG), improving code discovery, reuse, and the reliability of LLM-based coding. Yet existing code IR models remain largely text-centric and often overlook the visual and…
Beiji Lu
Synonymous codons encode the same amino acid yet are used non-randomly across genomes, a phenomenon with well-documented functional consequences for translation efficiency and mRNA stability. Whether the information embedded in synonymous codon choice is recoverable from the internal representations of in-dependently…
Yanan Shi, Chaoxiang Xie, Zhensu Sun, Yeheng Chen + 6 more
Large Language Models (LLMs) have achieved remarkable success in source code understanding, yet as software systems grow in scale, computational efficiency has become a critical bottleneck. Currently, these models rely on a text-based paradigm that treats source code as a linear sequence of tokens, which leads to a…
Yanshuo Chen, Yuming Zhang, Joshua Li, Boxue Tian + 1 more
Codon optimization involves selecting synonymous codons to match host-specific preferences. It is critical for heterologous expression but remains challenging due to the combinatorial design space. Under long-term evolutionary selection, natural coding sequences are near-optimal compromises between translational…
Andrea Gurioli, Federico Pennino, Maurizio Gabbrielli
Embedding-based code retrieval often suffers when encoders overfit to surface syntax. Prior work mitigates this by using LLMs to rephrase queries and corpora into a normalized style, but leaves two questions open: how much representational shift helps, and when is the per-query LLM call justified? We study a hierarchy…
Yuan Liu, Zhengmin Kong, Tao Huang, Yang Yang + 3 more
Nonintrusive load monitoring (NILM) is an effective approach for energy management that disaggregates the total power measured at the main power inlet into appliance-level power signals. NILM algorithms have achieved remarkable progress in recent years. However, accurately reconstructing appliance-level power signals…
Shiyi Du, Litian Liang, Jiayi Li, Carl Kingsford + 1 more
CodonMoE processes representations of codons, which are groups of three nucleotides in genetic sequences encoding amino acids. The input to CodonMoE consists of nucleotide representations with dynamic dimensionality, allowing it to accommodate input samples of varying sequence lengths. These inputs are reshaped into…
William Gilpin
Many large-scale pretrained models for genomic data directly adapt language architectures, treating the genome as a large body of text, with nucleotides acting as an alphabet and genes as words (; ; ; ; ; ). Genes thus may seem to be a more natural analogue to words in statistical learning frameworks. However, this…
Mei Lang, Xingyu Fang, Zhen Wang, Mingxuan Chen + 5 more
Although mRNA codon language models provide a generalizable framework for biological sequence design, effective CDS design requires both a learned sequence design space that captures biological constraints and context-configurable design preferences. Here we present CodonMamba, a codon language model framework for mRNA…
Authors not listed
Machine learning models are increasingly applied to heterogeneous materials datasets spanning different synthesis routes, measurement protocols, and structural classes. Although multi-task and representation-learning approaches are commonly used to improve predictive performance, the latent representations learned by…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…