13 papers · ranked by Valyu relevance
Yiwei Zhao, Qiushi Lin, Hongbo Kang, Guy E. Blelloch + 5 more
In this paper, we highlight a task-data orchestration abstraction that supports a range of distributed applications, including graph processing and key-value stores. Given a batch of tasks each requesting one or more data items, where both tasks and data are distributed across multiple machines, each task must get…
June Chen, Neal Xu, Gragas Huang, Bok Zhou + 1 more
The rapid growth of AI-generated content (AIGC) has enabled highquality creative production across diverse domains, yet existing systems face critical inefficiencies in throughput, resource utilization, and scalability under concurrent workloads. This paper introduces OnePiece, a large-scale distributed inference…
Nandy, Barenya Kumar, Rupesh Nasre
—Relational data, occurring in the real world, are often structured as graphs, which provide the logical abstraction required to make analytical derivations simpler. As graphs get larger, the irregular access patterns exhibited in most graph algorithms, hamper performance. This, along with NUMA and physical memory…
Benjamin Brock, Renato Golin
Many important applications across science, data analytics, and AI workloads depend on distributed matrix multiplication. Prior work has developed a large array of algorithms suitable for different problem sizes and partitionings including 1D, 2D, 1.5D, and 2.5D algorithms. A limitation of current work is that existing…
Jiang, Chenyu, Cai, Zhenkun + 8 more
Context parallelism has emerged as a key technique to support long-context training, a growing trend in generative AI for modern large models. However, existing context parallel methods rely on static parallelization configurations that overlook the dynamic nature of training data, specifically, the variability in…
Yijie Zhou, Shi Pu
Decentralized optimization has emerged as a critical paradigm for distributed learning, enabling scalable training while preserving data privacy through peer-to-peer collaboration. However, existing methods often suffer from communication bottlenecks due to frequent synchronization between nodes. We present Overlapping…
Shannon Kinkead, Jackson Wesley, Whit Schonbein, David DeBonis + 2 more
Performant all-to-all collective operations in MPI are critical to fast Fourier transforms, transposition, and machine learning applications. There are many existing implementations for all-to-all exchanges on emerging systems, with the achieved performance dependent on many factors, including message size, process…
Da Wei, Evangelia Kalyvianaki
This paper introduces Dodoor, an efficient randomized decentralized scheduler designed for task scheduling in modern data centers. Dodoor leverages advanced research on the weighted balls-into-bins model with -batched setting. Unlike other decentralized schedulers that rely on real-time probing of remote servers…
Piotr Luczynski, Tal Ben-Nun, Leighton Wilson, Brian Van Essen
Spatial Dataflow Architectures are an emerging hardware pattern in high-performance computing, whose mesh-connected fixed-memory processing elements are tailored for structured grid kernels with two-dimensional neighborhoods. However, practical multiphysics codes are often computed on unstructured grids, which induce…
David L. Cole, Jordan Jalving, Jonah Langlieb, Jesse D. Jenkins
1 Andlinger Center for Energy and Environment Princeton University, Princeton, NJ 08540, USA 2 Artificial Intelligence and Modeling Simulation Atomic Machines Inc., Emeryville, CA 94608, USA 3 Department of Computer Science Princeton University, Princeton, NJ 08540, USA 4 Department of Mechanical and Aerospace…
William Won, Kartik Lakhotia, Madhu Kumar, Sudarshan Srinivasan + 1 more
Distributed machine learning has become increasingly important due to the massive scale of large-scale generative models. Both model parameters and data are distributed across many compute devices, which requires frequent collective communications to synchronize activations and parameter updates. Such collective…
Mejgan Dedaj, Argyro Gailla, Theofanis Ioannou, Stamatia Kastrinaki + 8 more
This study assesses the scalability of process-based and thread-based schedulers for many-core shared-memory systems using a memory-intensive row-wise quick-sort workload on large three-dimensional tensors. The process-based evaluation considers bounded prolific, bounded collective, and three pipe-based…
Alan Malta Rodrigues, Douglas Thain
High-Throughput Computing (HTC) environments tailored for high-concurrency resource efficiency require sophisticated orchestration to manage petabyte-scale data across heterogeneous resources. A critical but often overlooked challenge is workflow composition: the strategic grouping of tasksets within a Directed Acyclic…