13 papers · ranked by Valyu relevance
Zhenkun Cai, Xiao Yan, Kaihao Ma, Yidi Wu + 4 more
'James Cheng' 'Teng Su' 'Fan Yu'] A good parallelization strategy can significantly improve the efficiency or reduce the cost for the distributed training of deep neural networks (DNNs). Recently, several methods have been proposed to find efficient parallelization strategies but they all optimize a single objective…
Sandeep U Mane, Pooja S. Lokare, Harsha R. Gaikwad
Ant Colony Optimization algorithm is a magnificent heuristics technique based on the behaviour of ants. Parallel computing is a mean to achieve the desired results in commensurable execution time. Parallelization of Ant Colony Optimization is utilized to solve large and complex problems. In this paper, review of…
Karame Mohammadiporshokooh, Steven R. Brandt, R. Tohid, Hartmut Kaiser
'Hartmut Kaiser'] Abstract. C++ Executors simplify the development of parallel algorithms by abstracting concurrency management across hardware architectures. They are designed to facilitate portability and uniformity of user-facing interfaces; however, in some cases they may lead to performance inefficiencies due to…
Jesper Larsson Träff
These lecture notes are designed to accompany an imaginary, virtual, undergraduate, one or two semester course on fundamentals of Parallel Computing as well as to serve as background and reference for graduate courses on High-Performance Computing, parallel algorithms and shared-memory multiprocessor programming. They…
Chuan-Chi Wang, Chun‐Yen Ho, Chia-Heng Tu, Shih‐Hao Hung
Particle Swarm Optimization (PSO) is a stochastic technique for solving the optimization problem. Attempts have been made to shorten the computation times of PSO based algorithms with massive threads on GPUs (graphic processing units), where thread groups are formed to calculate the information of particles and the…
Rabab Alkhalifa, Fatima Alkhomayes, Boushra Almazroua, Dana Alhaidan + 2 more
'Maryam Alothman' 'Jumana Almuhaidib'] The Traveling Salesman Problem (TSP) is a well-known NP-hard combinatorial optimization problem with wide-ranging applications in logistics, routing, and intelligent systems. Due to its factorial complexity, solving large-scale instances requires scalable and efficient algorithmic…
Hao Lin, Ke Wu, Jun Li, Wujun Li
Distributed learning is commonly used for training deep learning models, especially large models. In distributed learning, manual parallelism (MP) methods demand considerable human effort and have limited flexibility. Hence, automatic parallelism (AP) methods have recently been proposed for automating the parallel…
S.E. Tavares, Carmo P. Brás, A. L. Custódio, Vítor Duarte + 1 more
'Pedro D. Medeiros'] Direct Multisearch (DMS) is a Derivative-free Optimization class of algorithms suited for computing approximations to the complete Pareto front of a given Multiobjective Optimization problem. It has a well-supported convergence analysis and simple implementations present a good numerical…
François Pacaud, Michel Schanen, Sungho Shin, Daniel Adrian Maldonado + 1 more
'Daniel Adrian Maldonado' 'Mihai Anitescu'] We investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the…
Shuhei Watanabe, Neeratyoy Mallik, Edward M. Bergman, Frank Hutter
Zero-Cost Benchmarks Authors: ['Shuhei Watanabe' 'Neeratyoy Mallik' 'Edward M. Bergman' 'Frank Hutter'] Abstract While deep learning has celebrated many successes, its results often hinge on the meticulous selection of hyperparameters (HPs). However, the time-consuming nature of deep learning training makes HP…
Shiting Long, Gustavo Ramirez-Hidalgo, Andreas Frommer, Dirk Pleiter
Gauss-Seidel is a well-established iterative method for the solution of linear systems, and multicoloring has been widely used to increase parallelism in iterative solution techniques. Implementing multi-color Gauss-Seidel with conventional divide-and-conquer parallelization strategies, however, may be inefficient due…
Accorsi, Luca, Laganà, Demetrio + 6 more
We propose a parallel shared-memory schema to cooperatively optimize the solution of a Capacitated Vehicle Routing Problem instance with minimal synchronization effort and without the need for an explicit decomposition. To this end, we design FILO2 x as a single-trajectory parallel adaptation of the FILO2 algorithm…
Donald S. Ene, V.I.E Anireh
- Evaluating how well a whole system or set of subsystems performs is one of the primary objectives of performance testing. We can tell via performance assessment if the architecture implementation meets the design objectives. Performance evaluations of several parallel algorithms are compared in this study. Both…