13 papers · ranked by Valyu relevance
Fabio Cumbo, Jayadev Joshi, Daniel Blankenberg
Background – The integration of command-line tools into the Galaxy platform is crucial for making complex computational methods accessible to a broader audience and ensuring reproducible research. However, the manual development of tool wrappers (i.e., the XML files that define the user interface and execution logic in…
Charatvaraphan, Rujiphart, Chatchaiyadech, Bunradar + 12 more
Rujiphart Charatvaraphan\ , Bunradar Chatchaiyadech\ , Thitirat Sukijprasert\ , Chaiyong Ragkhitwetsagul\ , Morakot Choetkiertikul\ , Raula Gaikovina Kula† , Thanwadee Sunetnanta\ , Kenichi Matsumoto‡ \**Faculty of Information and Communication Technology, Mahidol University, Thailand †Graduate School of Information…
Mohayeminul Islam, Ajay Kumar Jha, May Mahmoud, Sarah Nadi
Library migration is the process of replacing a library with a similar one in a software project. Manual library migration is time consuming and error prone, as it requires developers to understand the Application Programming Interfaces (API) of both libraries, map equivalent APIs, and perform the necessary code…
Akira Tanaka, Yusuke Kawamoto
The reproducibility crisis in scientific research has received widespread recognition, thereby increasing the importance of meta-analyses that integrate statistical analyses from multiple studies. However, statistical methods often have ambiguous and implicit underlying assumptions, which can lead to their erroneous…
Yoann Marquer, Domenico Bianculli, Lionel C. Briand
Python is one of the most popular programming languages; as such, projects written in Python involve an increasing number of diverse security vulnerabilities. However, existing state-of-the-art analysis tools for Python only support a few vulnerability types. Hence, there is a need to detect a large variety of…
Meng Wang, Yue Ma, Majid Garoosi, Wenting Fan + 3 more
The rapid expansion of the Python ecosystem has fueled two distinct but converging threats: adversaries increasingly target the software supply chain via the Python Package Index (PyPI), while also building evasive, cross-platform malicious binaries compiled from source code written in Python. Current program analysis…
Jinhua Wang, Biswa Sengupta
Cross-language migration of large software systems is a persistent engineering challenge, particularly when the source codebase evolves rapidly. We present a methodology for LLM-assisted continuous code translation in which a large language model translates a production Rust codebase (648K LOC, 65 crates) into Python…
Zhenzhen Ren, Xinpeng Zhang, Zhenxing Qian, Yan Gao + 3 more
The integration of external tools is pivotal for empowering Large Language Model (LLM) agents with real-world capabilities. However, training these agents through direct, continuous interaction with diverse tools is often prohibitively expensive, slow, and introduces additional development and maintenance overhead. To…
Islem Bouzenia, Michael Pradel
Replacing hand-written code with library API calls is a common refactoring that can reduce code size, make code more idiomatic, and reuse well-tested implementations. Yet many library-adoption opportunities are hard to find automatically: the original code often does not mention the target library and may resemble the…
Zhouming Wu, Dakota Murray
Science advances not only through the accumulation of facts but also through the evolution of tools. Crucially, tools are rarely used in isolation. They form tool portfolios, combinations shaped by a discipline's workflows and analytical demands. Software, near-ubiquitous in modern research and traceable across the…
Mohamed Almukhtar, Anwar Ghammam, Hua Ming
As AI agents increasingly contribute to code development and maintenance, there is still limited empirical evidence on the quality and risk characteristics of their changes in real-world projects, particularly for refactoring-oriented contributions. It remains unclear how agent-authored refactoring edits affect…
Maaz, Muhammad, DeVoe, Liam + 4 more
Property-based testing (PBT) is a lightweight formal method, typically implemented as a randomized testing framework. Users specify the input domain for their test using combinators supplied by the PBT framework, and the expected properties or invariants as a unit-test function. The framework then searches for a…
Antonios Saravanos, John Pazarzis, Stavros Zervoudakis, Dongnanzi Zheng
Python libraries often need to maintain a stable public API even as internal implementations evolve, gain new backends, or depend on heavy optional libraries. In Python, where internal objects are easy to inspect and import, users can come to rely on "reachable internals" that were never intended to be public, making…