25 papers · ranked by Valyu relevance
Peng Ding, Rick Stevens
Third-party Python libraries introduce dependency management overhead, supply chain risk, and deployment friction in constrained environments. A natural question is how much of this ecosystem can be replicated using only Python's standard library -- and at what correctness and performance cost. We address this…
Antonios Saravanos, John Pazarzis, Stavros Zervoudakis, Dongnanzi Zheng
Python libraries often need to maintain a stable public API even as internal implementations evolve, gain new backends, or depend on heavy optional libraries. In Python, where internal objects are easy to inspect and import, users can come to rely on "reachable internals" that were never intended to be public, making…
Ragnar Bjornsson
We introduce ASH, a multi-scale, multi-theory modeling program for quantum mechanics (QM), molecular mechanics (MM), and hybrid calculations, written in the Python programming language. ASH is written in response to the increasingly diverse computational chemistry software landscape that features more QM and MM…
Mohayeminul Islam, Ajay Kumar Jha, May Mahmoud, Sarah Nadi
Library migration is the process of replacing a library with a similar one in a software project. Manual library migration is time consuming and error prone, as it requires developers to understand the Application Programming Interfaces (API) of both libraries, map equivalent APIs, and perform the necessary code…
Zhihan Liu, Aidan Cordero, Justin B. Kinney
Computationally designed DNA sequence libraries are essential components of massively parallel reporter assays (MPRAs), deep mutational scanning (DMS) experiments, and other multiplex assays of variant effect (MAVEs). They are also increasingly used in silico to analyze genomic AI models. Designing these libraries…
Islem Bouzenia, Michael Pradel
Replacing hand-written code with library API calls is a common refactoring that can reduce code size, make code more idiomatic, and reuse well-tested implementations. Yet many library-adoption opportunities are hard to find automatically: the original code often does not mention the target library and may resemble the…
Matthieu Vilain, Stéphane Aris-Brosou
The ever-growing amount of available biological data leads modern analysis to be performed on large datasets. Unfortunately, bioinformatics tools for preprocessing and analyzing data are not always designed to treat such large amounts of data efficiently. Notably, this is the case when encoding DNA and RNA sequences…
Han Zhang, John Jonides
We present PupEyes, an open-source Python package for preprocessing and visualizing pupil size and fixation data. PupEyes supports data collected from EyeLink and Tobii eye-trackers as well as any generic dataset that conforms to minimal formatting standards. Developed with current best practices, PupEyes provides a…
Endre Bakken Stovner, Max Ticó, Ester Muñoz del Campo, Joan Pallarès-Albanell + 3 more
Sequence interval algebra is key to modern bioinformatics. Pyranges v1 offers a Python Pandas-based interface to a comprehensive palette of Rust-powered operations (e.g., overlap, count, slice intervals), enabling the intuitive development of efficient pipelines for diverse sequence data, including gene annotations…
Zhihan Liu, Aidan Cordero, Justin B. Kinney
We have presented PoolParty, a Python package that replaces the ad hoc scripting commonly used to design oligonucleotide libraries with a declarative, composable, and flexible framework. We demonstrated PoolParty through three use cases: a DMS library for protein GB1; an MPRA library probing transcriptional regulatory…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Authors not listed
DNA-encoded libraries (DELs) have emerged as a powerful platform for screening ultra-large chemical spaces by leveraging DNA barcodes to tag and track individual small molecules. Recent work has shown that machine learning can enhance DEL based hit discovery by denoising sequencing artifacts and improving binder…
Matt Burridge, Zhen Ou, Katherine James, Gizem Buldum + 4 more
Advances in laboratory automation and AI-driven experimental design have increased the scale and complexity of data generated in synthetic biology. Whilst biofoundries provide significant resources and infrastructure to execute these experiments, most laboratories rely on isolated automated instruments and software…
Authors not listed
The analysis of molecular dynamics (MD) simulations is a critical but fragmented process, often requiring researchers to chain together multiple software tools and write bespoke scripts for routine structural and dynamic analyses. This workflow complexity creates a significant barrier to efficiency, standardization…
Jayadev Joshi, Fabio Cumbo, Daniel Blankenberg
R is widely used in statistical computing, data analysis, and bioinformatics. A key contributor to its success in bioinformatics and computational biology is the open-source project Bioconductor. As of its latest release (3.20), the Bioconductor community offers 2,289 software packages for biomedical research…
Fabian Woller, Lis Arend, Christian Fuchsberger, Markus List + 1 more
Not all state-of-the-art libraries offer the functionalities specific to our application, such as statistical test calculation on all feature pairs and pairwise missing data removal and parallelization options. Accordingly, different libraries were selected as competitors, depending on the statistical test, and manual…
Darshan Mandge, Anıl Tuncel, Aurélien Jaquier, Ilkan Kilic + 5 more
The diversity of labs, tools, and data formats in neuroscientific research has historically posed challenges for data sharing and collaboration. Researchers often needed to convert data between various formats and adapt their software to new environments, which could divert attention from their main research goals.…
Michael Antonov, Gábor Csárdi, Szabolcs Horvát, Kirill Müller + 7 more
Networks or graphs are widely used across the sciences to represent relationships of many kinds. The igraph ([https://igraph.org]()) software library supports graph construction, analysis, and visualisation, combining fast and robust performance with a low entry barrier. igraph pairs a fast core written in C with…
Authors not listed
With the rapid growth of chemical data and information, there is an increasing need for chemistry undergraduates to master Python tools for analyzing large chemical datasets and extracting key or feature information. Currently, more than 100,000 types of metal-organic frameworks (MOFs), as the material recently awarded…
Authors not listed
We present an open source collection of scripts and programs for the setup, management and evaluation of calculations with the Vienna ab-initio simulation package (VASP), called utils4VASP. It contains 20 independent Python scripts and Fortran programs, all with a unified and intuitive handling concept based on command…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Hasindu Gamaarachchi, Sasha Jenner, Hiruna Samarakoon, James M Ferguson + 1 more
Nanopore sequencing is a widespread and important method in genomics science. The raw electrical current signal data from a typical nanopore sequencing experiment are large and complex. This can be stored in 2 alternative file formats that are presently supported: POD5 is a signal data file format used by default on…
Hala Alia, Andrew Case, Irfan Ahmed
Modern software development relies heavily on third-party components from public repositories, expanding the software supply chain attack surface. In response to these growing risks, federal initiatives have advanced the Software Bill of Materials (SBOM) as a standardized mechanism for improving transparency by…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Authors not listed
Continuous manufacturing processes offer significant advantages over batch processes, including easier scalability, reduced costs, lower raw material and solvent consumption, and improved energy efficiency. A robust techno-economic assessment is therefore essential to evaluate and facilitate the adoption of such…