21 papers · ranked by Valyu relevance
Zhiming Zhang, Qingfu Zhu, Xianzhen Luo, Yixuan Wang + 2 more
Code translation aims to translate the code from its source language to the target language and is used in various software development scenarios. Recent developments in Large Language Models (LLMs) have showcased their capabilities in code translation, and parallel corpora play a crucial role in training models for…
Shahd Seddik, Fahd Seddik, Iman Saberi, Fatemeh H. Fard + 2 more
Large Language Models (LLMs) excel at code generation but struggle with complex problems. Retrieval-Augmented Generation (RAG) mitigates this issue by integrating external knowledge, yet retrieval models often miss relevant context, and generation models hallucinate with irrelevant data. We propose Programming…
Yufu Wang, He Jiang, Hao Lin, Peiyu Zou + 3 more
Large language models (LLMs) have shown great promise for automated code translation, yet existing approaches often rely on token-level statistical patterns rather than sufficient understanding of program semantics. As a result, translated programs may still contain logical and semantic errors. Although high-quality…
Zhengbin Zou, Tao Jiang, Yizheng Wang, Tiancheng Xue + 2 more
The increasing complexity of software systems has rendered code vulnerability detection a critical aspect of software security. While deep learning-based approaches have advanced this field, challenges such as coarse-grained function-level detection, scalability limitations, and constrained accuracy persist. Although…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Sun, Zhensu, Yang, Chengran + 8 more
—Large language models (LLMs) have shown exceptional performance in code generation and understanding tasks, yet their high computational costs hinder broader adoption. One important factor is the inherent verbosity of programming languages, such as unnecessary formatting elements and lengthy boilerplate code. This…
Sicong Liu, Yanxian Huang, Mingwei Liu, Ting Chen + 5 more
—Code generation tasks aim to automate the conversion of user requirements into executable code, significantly reducing manual development efforts and enhancing software productivity. The emergence of large language models (LLMs) has significantly advanced code generation, though their efficiency is still impacted by…
Zhuoyang Chen, Ruoqi Wang, Qiong Luo
Protein language models (pLMs) offer great potential for protein sequence analysis, yet the scarcity of labeled data often limits their effectiveness in fine-tuning. Data augmentation is a promising remedy, but systematic evaluation of augmentation strategies for protein sequences remains limited, and the conditions…
Adrien Mialland, Shuzo Fukunaga, Riku Katsuki, Yunfei Dong + 2 more
Data scarcity limits the characterization of protein fitness landscapes and the development of accurate variant effect prediction models. To address this challenge, we introduce fitness translocation, a data augmentation strategy that generates synthetic variants for a target protein by leveraging variant fitness data…
Christian Santamaria, Felipe Grijalva, Karen Rosero, José Vega-Sánchez + 3 more
Sound Event Localization and Detection (SELD) integrates Sound Event Detection (SED) and Direction-of-Arrival Estimation (DOAE) to recognize and localize sound events in various applications, including urban sound sensing, wildlife monitoring, and home surveillance. Recently, advancements in machine learning…
Ming-Feng Yeh, Ching-Chuan Luo, Cheng-Lin Lu, Nianbo Liu
Smart manufacturing relies on programmable logic controllers (PLCs) that translate sensor inputs into actuator commands. Generating PLC programs in legacy textual languages such as Mitsubishi FX-series Instruction List (IL) remains an expert-only task, and IL’s deprecation in IEC 61131-3 Edition 3.0 leaves it…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Tongcheng Geng, Muhammad Ahsan
Deep code models face security vulnerabilities through backdoor attacks. Previous approaches have primarily relied on single-trigger mechanisms, resulting in limited stealth and vulnerability to defense strategies. This paper proposes a novel hybrid backdoor attack method that combines function signature features and…
Niaz Bahar Chowdhury, August George, Sumit Purohit, Angela Cintolesi + 17 more
Genome-scale metabolic models (GEMs) are powerful tools for predicting cellular phenotypes and guiding microbial strain engineering, yet broad adoption remains challenging due to the computational expertise required. To overcome that, we present ChatGEM, an agentic platform that enables interactive GEM simulation…
Mingqi Wang, Yu Yang, Minna Gao, Jinliang Yuan + 2 more
Class imbalance remains a major obstacle to reliable network intrusion detection, particularly in Internet of Things (IoT) and sensor-network monitoring scenarios where rare attack categories are represented by only a small number of high-dimensional traffic samples. To improve minority-class augmentation, we propose…
Huanqiu Zhang, Israel Nelken, Tatyana Sharpee
Deciphering the neural code requires identifying its fundamental symbols or code-words. Neural activity is usually interpreted either as a rate code – based on average spike counts – or as a temporal code, which distinguishes patterns with identical counts. Yet, the symbols of the code remain undefined. Here we show…
E. G. Cooch, D. I. MacKenzie, J. A. Royle
Data augmentation is now a standard device across capture–recapture and occupancy analysis: adding a fixed number M of all-zero encounter histories replaces a model of unknown dimension with one of fixed dimension. Although M is often treated as a computational tuning choice, it also specifies a finite superpopulation…
GyeongTaek Choi, Seungho Jeon
WebAssembly is a low-level binary format originally designed to enable high-performance applications to run in web browsers. As WebAssembly is increasingly being ported to various environments, the security verification of WebAssembly execution environments is becoming more critical. While a wide variety of WebAssembly…
Aswathi Shiju, Samantha D. M. Arras, Allen G. Rodrigo, Anthony M. Poole + 1 more
In biology, changes to a DNA sequence can impact protein sequence but changes to protein sequences (phenotype) do not flow back into DNA (genotype). A system with bidirectional information flow (i.e., both translation and ‘reverse translation’) remains a theoretical possibility for an independent origin of life or an…
Authors not listed
This work provides a rigorous theoretical investigation of selective error correction strategies for variational quantum algorithms, with focus on understanding the interplay between error suppression, circuit trainability, and computational resource requirements. We develop a mathematical framework that characterizes…
Authors not listed
The axial ligand of heme is a key determinant of reactivity in heme-dependent enzymes, yet its systematic engineering remains challenging due to the limited chemical diversity of natural amino acids. Here, we demonstrate that replacing the native histidine axial ligand with a non-natural analog provides an effective…