22 papers · ranked by Valyu relevance
Vasilis Vouvoutsis, Constantinos Patsakis, Fran Casino
Malware research primarily studies the results, the methods, and the impact. Even from an offensive security perspective, what is examined is the method, not the development strategy of the offender. This study investigates the behavioral signatures and coding patterns embedded in the malware source code. By analyzing…
Yiwen Zhang, Wei Liu, Fazhong Jiang, Jiquan Ma + 4 more
Large Language Models of the Transformer architecture display great promise in automated code error detection based on their strength in processing sequential data. Nevertheless, their efficacy could be further improved by addressing the inherent weakness in handling structural code dependencies. In response to this…
Melih Peker, Ozcan Ozturk
Selecting a good set of optimization flags requires extensive effort and expert input. While most of the prior research considers using static, spatial, or dynamic features, some of the latest research directly applied deep neural networks to source code. We combined the static features, spatial features, and deep…
Jean-Charles Noirot Ferrand, Kyle Domico, Yohan Beugin, Patrick McDaniel
Open-source software (OSS) pipelines rely on automated static analysis tools to prevent the introduction of vulnerabilities in code. However, there is limited understanding of the efficacy of these tools across the OSS ecosystem over time. In this paper, we introduce a novel method to evaluate static application…
Xiaojian Liu, Yangyang Zhang, Chee Wei Tan, Wenyi Zhang
Code coverage-guided unit test generation (CGTG) and large language model-based test generation (LLMTG) are two principal approaches for the generation of unit tests. Each of these approaches has its inherent advantages and drawbacks. Tests generated by CGTG have been shown to exhibit high code coverage and high…
Huanqiu Zhang, Israel Nelken, Tatyana Sharpee
Deciphering the neural code requires identifying its fundamental symbols or code-words. Neural activity is usually interpreted either as a rate code – based on average spike counts – or as a temporal code, which distinguishes patterns with identical counts. Yet, the symbols of the code remain undefined. Here we show…
Yunkun Wang, Yue Zhang, Guochang Li, Zhi + 5 more
Large Language Models (LLMs) frequently generate buggy code with complex logic errors that are challenging to diagnose. While existing LLM-based self-repair approaches conduct intensive static semantic analysis or reply on superficial execution logs, they miss the in-depth runtime behaviors that often expose bug root…
Samuel W. Flint, Jigyasa Chauhan, Niloofar Mansoor, Bonita Sharif + 1 more
Modern programming languages, such as Python, support language features from several paradigms, such as object-oriented, procedural, and functional. Research has shown that code written in some paradigms can be harder to comprehend, but to date, no research has looked atwhich paradigm-specific language features impact…
Michael Pradel, Cristian Cadar, Islem Bouzenia
Numerous software analysis tools exist today, yet applying them to diverse open-source projects remains challenging due to environment setup, dependency resolution, and tool configuration. LLM-based agents offer a potential solution, yet no prior work has systematically studied their effectiveness on the specific task…
Tongcheng Geng, Muhammad Ahsan
Deep code models face security vulnerabilities through backdoor attacks. Previous approaches have primarily relied on single-trigger mechanisms, resulting in limited stealth and vulnerability to defense strategies. This paper proposes a novel hybrid backdoor attack method that combines function signature features and…
Authors not listed
We provide an overview of core molSimplify functionality and recent updates that enhance its capabilities for automated molecular and materials modeling. We describe the mol3D and atom3D classes, which store atomic and bonding information for a wide range of functions, including reading, modifying, and characterizing…
Elijah Zolduoarrati, Sherlock A. Licorish, Nigel Stanger
Developers frequently reuse Stack Overflow code snippets, yet the quality of these snippets remains unevenly understood, particularly across programming languages and geographic contexts. This study investigates code quality in Stack Overflow answers from contributors located in the United States, focusing on SQL…
Authors not listed
The analysis of molecular dynamics (MD) simulations is a critical but fragmented process, often requiring researchers to chain together multiple software tools and write bespoke scripts for routine structural and dynamic analyses. This workflow complexity creates a significant barrier to efficiency, standardization…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Yuxuan Zhang, Yiman Wang, Yang Tan, Yong Zhang
High-throughput assays generate diverse chromatin datasets that require flexible workflows and context-dependent parameter choices. Although large language models (LLMs) can assist analysis, unconstrained LLM-based execution often exhibits unstable behavior and limited reproducibility. We present ChromSkills, a curated…
Shuhong Huang, Ruben Portugues, James E. Fitzgerald
The ultimate goal of sensory coding is to extract and represent the cues required for adaptive motor output. This suggests that sensory codes and behavioral outcomes may align, and a variety of studies have argued that both biological and engineered sensory systems represent stimuli similarly when they play similar…
Pietro Braione, Giovanni Denaro, Luca Gugliemo, Elson Kurian + 2 more
Context. Since the eighties, the combination of program analysis techniques has been increasingly recognized as a promising approach to overcome the limitations of standalone methods. While individual techniques, based on either static or dynamic analysis, address important challenges in software dependability, their…
Jeremy Li, Alex Rubinsteyn, Sergey Feldman, Timothy O’Donnell + 18 more
Scientific computing has become a central component of modern scientific discovery. Yet many computational tools are developed by small, specialized teams under incentives that encourage the release of rapidly prototyped tooling without commensurate attention to engineering concerns, including performance and…
Virginia Iannibelli, Isabella Caranzano, Giovanni Birolo, Cesare Rollo + 5 more
Codon usage bias is a central record of mutation, selection, drift, and translational constraints, but it is usually treated separately from generalized Chargaff symmetry, the tendency for words and their reverse complements to occur at similar frequencies in long DNA sequences. Here we ask whether codon usage contains…
Authors not listed
Chemistry curricula often separate “wet” experimental work from “dry” computation, yet modern discovery increasingly demands both. This Perspective offers an instructor-ready roadmap to train “hybrid chemists” within existing courses. We distill recent advances in machine learning, automation, and real-time analytics…
Liyang Fei, Jovana Maksimovic, Alicia Oshlack
A “clone” encompasses a progenitor cell and its progeny cells. Tracking clonal composition as cells differentiate or evolve is useful in many fields. Various single-cell lineage tracing (clonal tracking) technologies use unique DNA barcodes that are passed from progenitor cells to their offspring. The barcode count for…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…