5 papers · ranked by Valyu relevance
Christos Tsolakis, Polykarpos Thomadakis, Nikos Chrisochoides
Handling the ever-increasing complexity of mesh generation codes along with the intricacies of newer hardware often results in codes that are both difficult to comprehend and maintain. Different facets of codes such as thread management and load balancing are often intertwined, resulting in efficient but highly complex…
Jing Hou, Guang Chen, Ruiqi Zhang, Zhijun Li + 2 more
'Changjun Jiang'] Abstract—The promotion of large-scale applications of reinforcement learning (RL) requires efficient training computation. While existing parallel RL frameworks encompass a variety of RL algorithms and parallelization techniques, the excessively burdensome communication frameworks hinder the…
Krzysztof Stuglik, Piotr Listkiewicz, Mateusz Kulczyk, Marcin Pietroń
'Marcin Pietroń'] Manual translation of the algorithms from sequential version to its parallel counterpart is time consuming and can be done only with the specific knowledge of hardware accelerator architecture, parallel programming or programming environment. The automation of this process makes porting the code much…
Denis Los, Igor Petushkov
Cores Authors: ['Denis Los' 'Igor Petushkov'] Abstract—Nowadays, latency-critical, high-performance applications are parallelized even on power-constrained client systems to improve performance. However, an important scenario of fine-grained tasking on simultaneous multithreading CPU cores in such systems has not been…
Wang, Haitian, Qin, Long
This paper presents an in-depth investigation into the high-performance parallel optimization of the Fish School Behaviour (FSB) algorithm on the Setonix supercomputing platform using the OpenMP framework. Given the increasing demand for enhanced computational capabilities for complex, large-scale calculations across…