14 papers · ranked by Valyu relevance
Cheng-Hsiang Chiu, Tsung‐Wei Huang, Zizheng Guo, Yibo Lin
Pipeline is a fundamental parallel programming pattern. Mainstream pipeline programming frameworks count on data abstractions to perform pipeline scheduling. This design is convenient for data-centric pipeline applications but inefficient for algorithms that only exploit task parallelism in pipeline. As a result, we…
Houming Wu, Ling Chen, Wenjie Yu
Large Models Training Authors: ['Houming Wu' 'Ling Chen' 'Wenjie Yu'] With the increasing scale of models, the need for efficient distributed training has become increasingly urgent. Recently, many synchronous pipeline parallelism approaches have been proposed to improve training throughput. However, these approaches…
Jesper Larsson Träff
These lecture notes are designed to accompany an imaginary, virtual, undergraduate, one or two semester course on fundamentals of Parallel Computing as well as to serve as background and reference for graduate courses on High-Performance Computing, parallel algorithms and shared-memory multiprocessor programming. They…
Byungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim + 10 more
Graph Pipeline Parallelism Authors: ['Byungsoo Jeon' 'Mengdi Wu' 'Shiyi Cao' 'Sunghyun Kim' 'Sunghyun Park' 'Neeraj Aggarwal' 'Colin Unger' 'Daiyaan Arfeen' 'Peiyuan Liao' 'Xupeng Miao' 'Mohammad Reza Alizadeh' 'Gregory R. Ganger' 'Tianqi Chen' 'Zhihao Jia'] Deep neural networks (DNNs) continue to grow rapidly in size…
Branden Butler, Sixing Yu, Arya Mazaheri, Ali Jannesari
Speculation Authors: ['Branden Butler' 'Sixing Yu' 'Arya Mazaheri' 'Ali Jannesari'] Abstract—Inference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from CPU speculative execution. These techniques…
A. Dutta, Nabendu Chaki, Rajat K. De
Removed Staleness Authors: ['A. Dutta' 'Nabendu Chaki' 'Rajat K. De'] DNN training is extremely time-consuming, necessitating efficient multi-accelerator parallelization, where a single iteration of training is split over the available accelerators. Current approaches often parallelize training by using intra-batch…
Ye Tian, Zhen Jia, Ziyue Luo, Yida Wang + 1 more
Diffusion models have emerged as dominant performers for image generation. To support training large diffusion models, this paper studies pipeline parallel training of diffusion models and proposes DiffusionPipe, a synchronous pipeline training system that advocates innovative pipeline bubble filling technique…
B. T. Fleming, K. Knoepfel, Meifeng Lin, X. Qian + 5 more
'Brett Viren' 'H. Wei' 'Shinjae Yoo' 'Haiwang Yu'] DUNE, like other HEP experiments, faces a challenge related to matching execution patterns of our production simulation and data processing software to the limitations imposed by modern highperformance computing facilities. In order to efficiently exploit these new…
Jihu Guo, T.P. Ma, Wei Gao, Peng Sun + 4 more
Pipeline parallelism is widely used to train large language models (LLMs). However, increasing heterogeneity in model architectures exacerbates pipeline bubbles, thereby reducing training efficiency. Existing approaches overlook the cooptimization of model partition, model placement, and workload scheduling, resulting…
Temitayo Adefemi
—Parallelization has become a cornerstone of modern computing, influencing everything from high-performance supercomputers to everyday mobile devices. This paper presents a comprehensive guide on the fundamentals of parallelization that every computer scientist should know, beginning with a historical perspective that…
Donald S. Ene, V.I.E Anireh
- Evaluating how well a whole system or set of subsystems performs is one of the primary objectives of performance testing. We can tell via performance assessment if the architecture implementation meets the design objectives. Performance evaluations of several parallel algorithms are compared in this study. Both…
Denis Los, Igor Petushkov
Cores Authors: ['Denis Los' 'Igor Petushkov'] Abstract—Nowadays, latency-critical, high-performance applications are parallelized even on power-constrained client systems to improve performance. However, an important scenario of fine-grained tasking on simultaneous multithreading CPU cores in such systems has not been…
Rajendra Purohit, K. R. Chowdhary, Sunıl Dutt Purohıt
—Arrival of multicore systems has enforced a new scenario in computing, the parallel and distributed algorithms are fast replacing the older sequential algorithms, with many challenges of these techniques. The distributed algorithms provide distributed processing using distributed file systems and processing units…
Patrick Mukala
| Article Info | ABSTRACT | | --- | --- | | | A myriad of applications ranging from engineering and scientific | | | simulations, image and signal processing as well as high-sensitive data | | | retrieval demand high processing power reaching up to teraflops for their | | | efficient execution. While a standard serial…