12 papers · ranked by Valyu relevance
Stephen Mell, David Mell, Konstantinos Kallas, Steve Zdancewic + 1 more
Compound AI applications, which compose calls to ML models using a general-purpose programming language like Python, are widely used for a variety of user-facing tasks, from software engineering to enterprise automation, making their end-to-end latency a critical bottleneck. In contrast to traditional applications…
Tsai, Shin-Rong, Schive, Hsi-Yu + 2 more
In the exascale computing era, handling and analyzing massive datasets have become extremely challenging. In situ analysis, which processes data during simulation runtime and bypasses costly intermediate I/O steps, offers a promising solution. We present libyt (https://github.com/yt-project/libyt), an open-source C…
Xuesong Geng, Yunwei Cui, Lingang Zhang, Liangliang Ji
We present $λ$PIC, a Python-based electromagnetic particle-in-cell framework built around a callback-centric architecture. Existing PIC codes typically tie high performance to static, pre-compiled timestep loops, hindering implementation of custom physics, diagnostics, or output logic. $λ$PIC breaks this coupling by…
Lizhe Chen
This report describes Infernux, an open-source game engine that pairs a C++17/Vulkan real-time core with a Python production layer connected through a single pybind11 boundary. To close the throughput gap between Python scripting and native-code engines, Infernux combines two established techniques - batch-oriented…
Marc Becker, Bernd Bischl
Many algorithms in statistics and machine learning can be parallelized in an asynchronous manner where workers need to communicate through shared state rather than execute independent tasks dispatched by a central controller. Especially in modern hyperparameter optimization and parallel black-box optimization with…
Shiqi Cheng, Evelyne Ringoot, Rabab Alomairy, Alan Edelman
AI coding agents have quickly become omnipresent in software engineering. Their serial performance, both in terms of accuracy and speed, has been extensively covered. However, recent initial results suggest their parallel programming capabilities lag behind serial programming capabilities. This paper presents a…
Haymo Kutschbach
This work introduces a self-optimizing virtual processor (VP) for numerical array programs that shifts parallelization from a manual developer task to a cooperative, agent-like runtime mechanism. Instead of relying on centralized task-graph scheduling, static compiler optimization, or explicitly annotated parallel…
Shiting Long, Gustavo Ramirez-Hidalgo, Andreas Frommer, Dirk Pleiter
Gauss-Seidel is a well-established iterative method for the solution of linear systems, and multicoloring has been widely used to increase parallelism in iterative solution techniques. Implementing multi-color Gauss-Seidel with conventional divide-and-conquer parallelization strategies, however, may be inefficient due…
Yanda Tao, Pedro F. Silvestre, Marcel Wagenländer, Peter Pietzuch
The scale of LLM training jobs requires parallelization planning over large GPU clusters. Due to different GPU types and interconnects added over time, these GPU clusters are increasingly heterogeneous. Automatic LLM parallelizers can search for parallelization plans but face an exploding search space with…
Hala ElAarag, Anas Gamal Aly
Parallel and Distributed Computing (PDC) is a critical yet conceptually challenging area of the undergraduate computer science curriculum. While students often encounter these concepts in theory, few gain exposure to experience in real high-performance computing (HPC) environments. Research shows that when students are…
Paulo Henrique Leme Ramalho, Dennis Alves Pedersen, Fábio Andrijauskas
The complexity of biomolecular simulations has substantially increased the demand for High-Performance Computing (HPC) infrastructures, particularly in molecular dynamics and coarse-grained modeling. This work presents a systematic performance and scalability analysis of the LAMMPS simulator for coarse-grained…
Atharva Chougule, Alexander J Root, Rubens Lacouture, Bobby Yan + 2 more
Sparse tensor algebra is challenging to efficiently parallelize due to the irregular, data-dependent, and potentially skewed structure of sparse computation. We propose the first partitioning algorithm that provably load balances the computation of any sparse tensor algebra expression across parallel execution units.…