24 papers · ranked by Valyu relevance
Antonios Saravanos, John Pazarzis, Stavros Zervoudakis, Dongnanzi Zheng
Python libraries often need to maintain a stable public API even as internal implementations evolve, gain new backends, or depend on heavy optional libraries. In Python, where internal objects are easy to inspect and import, users can come to rely on "reachable internals" that were never intended to be public, making…
Zeyi Zhang, Carlos Mora Perez, Patrick Kwon, Martin Head‐Gordon + 1 more
PARSEC.py is a Python-based real-space Kohn-Sham density functional theory (real-space KS-DFT) framework designed to provide a user- and developer-friendly platform for first-principles electronic-structure simulations. Discretization on real-space grids eliminates basis-set approximations, while enabling systematic…
Robert Wolff, Alessia Polito, Alessio Paolo Buccino, Michela Chiappalone + 1 more
Title: Summary High-density multi-electrode arrays enable the recording of in vitro neuronal activity with exceptional spatial and temporal resolution. Here, we describe a protocol for analyzing these extensive datasets by using two complementary tools. The nicespike tool implements a full electrophysiological data…
Stephen R Piccolo, Harlan P Stevens
Using OpenAI’s Chat Completions API, we evaluated the ability to translate the example solutions and test code for the Python exercises to other programming languages: C++, Rust, Julia, and JavaScript. When invoking the API, we used version “gpt-4-0314” of the model and the default temperature setting of 0.7. For each…
Samuel W. Flint, Jigyasa Chauhan, Niloofar Mansoor, Bonita Sharif + 1 more
Modern programming languages, such as Python, support language features from several paradigms, such as object-oriented, procedural, and functional. Research has shown that code written in some paradigms can be harder to comprehend, but to date, no research has looked atwhich paradigm-specific language features impact…
Authors not listed
The analysis of molecular dynamics (MD) simulations is a critical but fragmented process, often requiring researchers to chain together multiple software tools and write bespoke scripts for routine structural and dynamic analyses. This workflow complexity creates a significant barrier to efficiency, standardization…
Endre Bakken Stovner, Max Ticó, Ester Muñoz del Campo, Joan Pallarès-Albanell + 3 more
Sequence interval algebra is key to modern bioinformatics. Pyranges v1 offers a Python Pandas-based interface to a comprehensive palette of Rust-powered operations (e.g., overlap, count, slice intervals), enabling the intuitive development of efficient pipelines for diverse sequence data, including gene annotations…
Han Zhang, John Jonides
We present PupEyes, an open-source Python package for preprocessing and visualizing pupil size and fixation data. PupEyes supports data collected from EyeLink and Tobii eye-trackers as well as any generic dataset that conforms to minimal formatting standards. Developed with current best practices, PupEyes provides a…
Authors not listed
With the rapid growth of chemical data and information, there is an increasing need for chemistry undergraduates to master Python tools for analyzing large chemical datasets and extracting key or feature information. Currently, more than 100,000 types of metal-organic frameworks (MOFs), as the material recently awarded…
Tsai, Shin-Rong, Schive, Hsi-Yu + 2 more
In the exascale computing era, handling and analyzing massive datasets have become extremely challenging. In situ analysis, which processes data during simulation runtime and bypasses costly intermediate I/O steps, offers a promising solution. We present libyt (https://github.com/yt-project/libyt), an open-source C…
Authors not listed
We present an open source collection of scripts and programs for the setup, management and evaluation of calculations with the Vienna ab-initio simulation package (VASP), called utils4VASP. It contains 20 independent Python scripts and Fortran programs, all with a unified and intuitive handling concept based on command…
Lior Pachter
The edgeR Bioconductor package is one of the most widely used tools for differential expression analysis of count-based genomics data. Despite its popularity, the R-only implementation limits its integration with the Python-centric ecosystem that has become dominant in single-cell genomics. We present edgePython, a…
Sicong Liu, Yanxian Huang, Mingwei Liu, Ting Chen + 5 more
—Code generation tasks aim to automate the conversion of user requirements into executable code, significantly reducing manual development efforts and enhancing software productivity. The emergence of large language models (LLMs) has significantly advanced code generation, though their efficiency is still impacted by…
Matthieu Vilain, Stéphane Aris-Brosou
The ever-growing amount of available biological data leads modern analysis to be performed on large datasets. Unfortunately, bioinformatics tools for preprocessing and analyzing data are not always designed to treat such large amounts of data efficiently. Notably, this is the case when encoding DNA and RNA sequences…
Shuo Yang, Yuting Liu, Jie Mi, Li‐Jen Chang + 1 more
According to the traditional Chinese veterinary medicine (TCVM) theory, Yin deficiency contributes to the acute onset of eclampsia. This condition arises primarily from the substantial loss of blood and nutrients to the foetus during pregnancy, leading to insufficient nutrition and blood supply. Consequently, the…
Authors not listed
TurtleMol is an open-source Python package that aims to help users generate large, complex molec- ular systems. In the current version, users can generate systems by filling volumes defined by basic geometric shapes (e.g. cube, sphere), or by shapes of arbitrary gemoetries defined meshes created in other software (such…
Yajushi Khurana, Keisuke Ishihara
Three-dimensional biological morphologies encode functional and physiological state, yet the directional, orientational, and topological properties of these shapes are rarely captured by morphometric tools available for bioimage analysis. Minkowski tensors are mathematically rigorous tensor-valued measures that encode…
Mohamed Almukhtar, Anwar Ghammam, Hua Ming
As AI agents increasingly contribute to code development and maintenance, there is still limited empirical evidence on the quality and risk characteristics of their changes in real-world projects, particularly for refactoring-oriented contributions. It remains unclear how agent-authored refactoring edits affect…
Authors not listed
Computational methods for predictive modeling have been increasingly utilized in the early stages of drug discovery to supplement high-throughput screening. The advent of highly efficient and complex machine learning architectures necessitates new methods of collating the plethora of topological, geometrical, and…
Michael Antonov, Gábor Csárdi, Szabolcs Horvát, Kirill Müller + 7 more
Networks or graphs are widely used across the sciences to represent relationships of many kinds. The igraph ([https://igraph.org]()) software library supports graph construction, analysis, and visualisation, combining fast and robust performance with a low entry barrier. igraph pairs a fast core written in C with…
Jinhua Wang, Biswa Sengupta
Cross-language migration of large software systems is a persistent engineering challenge, particularly when the source codebase evolves rapidly. We present a methodology for LLM-assisted continuous code translation in which a large language model translates a production Rust codebase (648K LOC, 65 crates) into Python…
Jane P. Wesson, Mark A. O'Dea, Andrew J. Currie, Cathy M. Shilton + 2 more
Sunshinevirus was originally isolated in 2008 from a collection of Australian pythons suffering an outbreak of neuro-respiratory disease. Sunshinevirus infection cases, including chronic asymptomatic infections, have subsequently been detected throughout Australia by PCR testing. To investigate the association between…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Kenji Gerhardt, Shujun Ou
The rapid expansion of high-quality, nearly complete eukaryotic genomes demands computational accelerations of existing bioinformatic infrastructure. TIR-Learner has been widely used for the de novo identification of Terminal Inverted Repeat (TIR) transposons, but suffers from slow runtime and a large memory footprint.…