26 papers · ranked by Valyu relevance
Panchal, Deven
—Generative Agentic AI systems are emerging as a powerful paradigm for automating complex, multi-step tasks. However, many existing frameworks for building these systems introduce significant complexity, a steep learning curve, and substantial boilerplate code, hindering rapid prototyping and deployment. This paper…
Asif Zaman, Kallol Naha, Khalid Belhajjame, Hasan M. Jamil
Scientific workflows encode valuable domain expertise and computational methodologies. Yet studies consistently show that a significant proportion of published workflows suffer from decay over time. This problem is particularly acute for legacy workflow systems like Taverna, where discontinued services, obsolete…
Daniel Pearson, Sidney Shapiro, Emiliano Sebastian Gonzalez Venegas, Sanad Al-Khatib + 1 more
This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative AI systems in business processes. Rather than treating LangGraph, a low-level orchestration framework for stateful agents, as a model-quality benchmark target, we present three executable recipes -- SQL…
Mohamed Salim Nasser Al Hinai, Zakira Naureen, Syed Abdullah Gilani
Since, Bioinformatic research is getting attraction that can be seen with sudden increase in development of tools as well as publications. However, there are challenges in bioinformatics which are making the results difficult to obtain while facing reproducibility, scalability, and accessibility of computational…
Leonardo Pelonero, Fabio Vitello, Sciacca Eva, Mauro Imbrosciano + 2 more
—In recent years, the monitoring and study of natural hazards have gained significant attention, particularly due to climate change, which exacerbates incidents like floods, droughts, storm surges, and landslides. Together with the constant risk of earthquakes, these climate-induced events highlight the critical…
Thuan Ha, Kwabena Abrefa Nketia, Hansanee Fernando, Sarah van Steenbergen + 2 more
Accurate field boundary delineation is critical for accurate modelling on crop yields and for precision agriculture (PA), enabling site-specific management to optimize resource use and crop productivity. Traditional boundary mapping methods, such as manual digitization and semi-automated extraction from farm machinery…
Zhengxue Zhou, Satheeshkumar Veeramani, Francisco Munguia-Galeano, Hatem Fakhruldeen + 1 more
Self-driving labs (SDLs) combine robotic automation with artificial intelligence (AI) to allow autonomous, high-throughput experimentation. However, robot manipulation in most SDL workflows operates in an open-loop manner, lacking real-time error detection and error correction. This can reduce reliability and overall…
Damien J. Mannion, Maria del Mar Quiroga, Jacob M. Paul, Marta I. Garrido
The processing of neuroimaging data typically involves a complicated set of operations, which often require different software packages and have intensive computational and storage demands. Although there are many options available for the neuroimaging researcher to establish their preferred set of processing…
Sidney Shapiro, Daniel Pearson, Emiliano Sebastian Gonzalez Venegas
Spreadsheet-heavy analytical work remains common in business analytics, operations reporting, and applied research, yet workbooks that grow through formulas, manual edits, and copy-paste refresh are difficult to audit, reproduce, and govern at scale. When tabular work requires repeatability, validation, version…
Jon Marcos-Mercadé, Unai Lopez-Novoa, Mikel Egaña Aranguren
Given the increasing adoption of AI solutions in professional environments, it is necessary for developers to be able to make informed decisions about the current tool landscape. This work empirically evaluates various MLOps (Machine Learning Operations) tools to facilitate the management of the ML model lifecycle…
Matt Burridge, Zhen Ou, Katherine James, Gizem Buldum + 4 more
Advances in laboratory automation and AI-driven experimental design have increased the scale and complexity of data generated in synthetic biology. Whilst biofoundries provide significant resources and infrastructure to execute these experiments, most laboratories rely on isolated automated instruments and software…
Shixiang Wang
Bioinformatics analyses depend on workflow engines to coordinate dozens of computational tools across complex dependency chains. The most widely adopted engines—Snakemake, Nextflow, the Common Workflow Language (CWL), and the Workflow Description Language (WDL)—run on interpreted or just-in-time (JIT) compiled language…
Authors not listed
The analysis of molecular dynamics (MD) simulations is a critical but fragmented process, often requiring researchers to chain together multiple software tools and write bespoke scripts for routine structural and dynamic analyses. This workflow complexity creates a significant barrier to efficiency, standardization…
Authors not listed
With the rapid growth of chemical data and information, there is an increasing need for chemistry undergraduates to master Python tools for analyzing large chemical datasets and extracting key or feature information. Currently, more than 100,000 types of metal-organic frameworks (MOFs), as the material recently awarded…
Trina De, Vardan Andriasyan, Artur Yakimovich
Virological plaque assays are the primary method for quantifying infectious particles in a suspension, achieved by incubating a serial dilution of the virus with a monolayer of indicator cells. Existing software tools for quantification of plaque assay images lack modularity, show measurements disagreement or are…
Julie R. Pivin-Bachler, Egon L. van den Broek
Title: Summary Machine learning struggles with imbalanced data. Although several mitigation approaches exist, their application depends on the extent of imbalance. To determine the latter, a protocol was developed. Across 428 synthetic and 70 real datasets, 8 imbalance measures were benchmarked and evaluated using…
Mingda Zhang, Haoran Luo, Tiesunlong Shen, Qika Lin + 3 more
In recent years, a variety of powerful agentic workflows have been applied to solve a wide range of human problems. However, existing workflow orchestration still faces key challenges, including high manual cost, reliance on specific operators/LLMs, and sparse reward signals. To address these challenges, we propose…
Authors not listed
We present an open source collection of scripts and programs for the setup, management and evaluation of calculations with the Vienna ab-initio simulation package (VASP), called utils4VASP. It contains 20 independent Python scripts and Fortran programs, all with a unified and intuitive handling concept based on command…
Authors not listed
Self-driving laboratories (SDLs) promise accelerated scientific discovery and product development by closing the loop between robotic execution and AI/ML-driven decision making. In practice, however, SDL orchestration remains fragmented; workflows are typically encoded as laboratory-specific scripts or bespoke…
Diego Fernández, Julián García-Vinuesa, Diego Alvarez-Saravia, Michelle Soto-García + 8 more
Biomolecular dataset construction often requires coordinated retrieval from heterogeneous repositories, identifier mapping, cross-reference enrichment, source-specific parsing, and provenance recording. These operations are frequently implemented through project-specific scripts, making acquisition procedures difficult…
James A. London, Abhishek K. Singh, Teague C. Svendsen, Naciye Esma Tirtom + 2 more
Magnetic tweezers are a popular biophysical instrument for manipulating and measuring single molecules. Most groups rely on custom-built setups tailored to specific experiments, making it challenging to implement and share software. Typically, image acquisition and hardware control are automated via LabVIEW, while…
Scott Huberty, James Desjardins, Tyler Collins, Mayada Elsabbagh + 1 more
EEG recordings are typically long and contain large amounts of data, making manual cleaning a time-consuming and error-prone task. Automated preprocessing pipelines can facilitate the efficient and objective extraction of artifacts, enabling standardized and reproducible analyses. However, automated preprocessing…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Authors not listed
Data-driven approaches offer great potential for accelerating ab initio electronic structure calculations of molecules and materials but their transferability is often limited due to the vast amount of data needed for training, including when addressing the need to fine-tune universal models for each specific system to…
Alexander Alsalihi, Robert M. Flight, Hunter N. B. Moseley
The recount3 online resource provides tens of thousands of uniformly processed RNA-seq samples across human and mouse from major sequencing repositories like the Sequence Read Archive. While access to these datasets has traditionally been centered in the R/Bioconductor ecosystem, the growing prominence of Python in…