24 papers · ranked by Valyu relevance
Prathamesh Devadiga
Traditional auto-parallelizing compilers, reliant on rigid heuristics, struggle with the complexity of modern heterogeneous systems. This paper presents a comprehensive evaluation of small ( 1B parameter) Language Model (LLM)-driven compiler auto-parallelization. We evaluate three models—gemma3, llama3.2, and…
Erel Kaplan, Tomer Bitan, Lian Ghrayeb, Le Chen + 3 more
Parallel programming is central to modern highperformance computing, but producing parallel implementations that are both correct and fast remains arduous. In practice, developers face two recurring needs: introducing parallelism into serial kernels to exploit CPUs and GPUs, and migrating existing parallel code between…
Martin Petr, Isabel M. Pötzsch, Fernando Racimo
Simulation-based inference methods such as Approximate Bayesian Computation (ABC) are a popular class of techniques in evolutionary biology and population genetics. These methods are particularly useful for fitting complex models, as they can bypass the need to compute an exact likelihood function, instead relying on…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
Authors not listed
Recent advances in machine learning force fields (MLFF) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference datasets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and…
Bowen Zhou, Jinrui Jia, Wenhao He, Yong Zhang + 1 more
—The Mixture of Experts (MoE) models are emerging as the latest paradigm for Large Language Models (LLMs). However, due to memory constraints, MoE models with billions or even trillions of parameters can only be deployed in multi-GPU or even multi-node & multi-GPU based serving systems. Thus, communication has became a…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
Yacine Hakimi, Riyadh Baghdadi
—As the demand for computational power grows, optimizing code through compilers becomes increasingly crucial. In this context, we focus on fully automatic code optimization techniques that automate the process of selecting and applying code transformations for better performance without manual intervention.…
Stephen Mell, David Mell, Konstantinos Kallas, Steve Zdancewic + 1 more
Compound AI applications, which compose calls to ML models using a general-purpose programming language like Python, are widely used for a variety of user-facing tasks, from software engineering to enterprise automation, making their end-to-end latency a critical bottleneck. In contrast to traditional applications…
Haymo Kutschbach
This work introduces a self-optimizing virtual processor (VP) for numerical array programs that shifts parallelization from a manual developer task to a cooperative, agent-like runtime mechanism. Instead of relying on centralized task-graph scheduling, static compiler optimization, or explicitly annotated parallel…
Mateusz Gruzewski, Marek Palkowski, Ramon Antonio Rodriges Zalipynis
In this article, we present an efficient and concise OpenMP implementation of the Nussinov RNA folding algorithm, a well-known representative of non-serial polyadic dynamic programming (NPDP). Our goal is to develop an optimized implementation that can serve as a template for related dynamic programming applications.…
Kenshin Obi, Takumi Onozawa, Fujimoto, Hiroshi + 1 more
—In recent years, autonomous vehicles have attracted attention as one of the solutions to various social problems. However, autonomous driving software requires real-time performance as it considers a variety of functions and complex environments. Therefore, this paper proposes a parallelization method for autonomous…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Marco Savioli, Paolo Calligari, Ugo Locatelli, Gianfranco Bocchinfuso
We introduce GROMODEX, a novel tool designed to optimise GROMACS molecular dynamics (MD) simulations using a structured Design of Experiments (DoE) approach. GROMACS, though efficient, requires extensive tuning of parameters to perform optimally on different hardware and molecular systems. Manual tuning is tedious and…
Noam Teyssier, Alexander Dobin
Single-cell genomics is rapidly scaling toward billion-cell atlases, but computational analysis has become a critical bottleneck. Processing multiplexed datasets with existing tools requires substantial computational resources and runtime that become prohibitive at scale. Here we present cyto, an ultra highthroughput…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
Background The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated…
Christoph Stelz, Lukas Hübner, Alexandros Stamatakis, Russell Schwartz
Next to disseminating results, scientific publications also aim to convince the reader of their validity (). While reproducibility is crucial for validating scientific claims (), practical attempts to reproduce computational findings frequently fail, e.g. in climate and weather modeling (), power grid analysis (), or…
Authors not listed
Bridging AI and self-driving laboratories, we introduce the first fully-automated, closed-loop molecular discovery cycle, exemplified by the identification of novel JAK inhibitors. With minimal human intervention, we combined AI-driven molecular design and retrosynthesis with IBM’s synthesis automation system RoboRXN…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Authors not listed
Predicting how chemical modifications affect drug binding is central to rational drug design. Free Energy Perturbation (FEP) calculations provide accurate estimates of these binding affinity changes, but existing methods often require substantial computational resources and expert knowledge. Here we present QligFEP…
Authors not listed
The analysis of molecular dynamics (MD) simulations is a critical but fragmented process, often requiring researchers to chain together multiple software tools and write bespoke scripts for routine structural and dynamic analyses. This workflow complexity creates a significant barrier to efficiency, standardization…
Authors not listed
Agentic artificial intelligence (AI) is poised to redefine how science is conducted, automating not just data analysis but the entire research lifecycle, from hypothesis generation to validation. Yet most current AI agents remain domain-bound, tailored to specific applications such as materials synthesis or quantum…