14 papers · ranked by Valyu relevance
Rishi Sharma, Shreyansh Kulshreshtha, Manas Thakur
With the advent of multi-core systems, GPUs and FPGAs, loop parallelization has become a promising way to speedup program execution. In order to stay up with time, various performance-oriented programming languages provide a multitude of constructs to allow programmers to write parallelizable loops. Correspondingly…
Donald S. Ene, V.I.E Anireh
- Evaluating how well a whole system or set of subsystems performs is one of the primary objectives of performance testing. We can tell via performance assessment if the architecture implementation meets the design objectives. Performance evaluations of several parallel algorithms are compared in this study. Both…
Rajendra Purohit, K. R. Chowdhary, Sunıl Dutt Purohıt
—Arrival of multicore systems has enforced a new scenario in computing, the parallel and distributed algorithms are fast replacing the older sequential algorithms, with many challenges of these techniques. The distributed algorithms provide distributed processing using distributed file systems and processing units…
Alexandru Calotoiu, Tal Ben‐Nun, Grzegorz Kwaśniewski, Johannes de Fine Licht + 3 more
'Johannes de Fine Licht' 'Timo Schneider' 'Philipp Schaad' 'Torsten Hoefler'] C is the lingua franca of programming and almost any device can be programmed using C. However, programming modern heterogeneous architectures such as multi-core CPUs and GPUs requires explicitly expressing parallelism as well as…
Temitayo Adefemi
—Parallelization has become a cornerstone of modern computing, influencing everything from high-performance supercomputers to everyday mobile devices. This paper presents a comprehensive guide on the fundamentals of parallelization that every computer scientist should know, beginning with a historical perspective that…
Weidong Wang, Haoran Zhu
Efficient Automated OpenMP Parallelization Authors: ['Weidong Wang' 'Haoran Zhu'] In advancing parallel programming, particularly with OpenMP, the shift towards NLP-based methods marks a significant innovation beyond traditional S2S tools like Autopar and Cetus. These NLP approaches train on extensive datasets of…
Urmila Shrawankar, Mayuri Joshi
— In multi-core systems, various factors like inter-process communication, dependency, resource sharing and scheduling, level of parallelism, synchronization, number of available cores etc. influence the extent of possible High Performance Computing parallelization. These parameters if not managed to the root level…
Amirhossein Shahbazinia, Saber Salehkaleybar, Matin Hashemi
—One of the key objectives in many fields in machine learning is to discover causal relationships among a set of variables from observational data. In linear non-Gaussian acyclic models (LiNGAM), it can be shown that the true underlying causal structure can be identified uniquely from merely observational data.…
Linus Zwaka
and MPI Authors: ['Linus Zwaka'] Abstract. Sequence alignment is a cornerstone of bioinformatics, widely used to identify similarities between DNA, RNA, and protein sequences and studying evolutionary relationships and functional properties. The Needleman-Wunsch algorithm remains a robust and accurate method for global…
Krzysztof Stuglik, Piotr Listkiewicz, Mateusz Kulczyk, Marcin Pietroń
'Marcin Pietroń'] Manual translation of the algorithms from sequential version to its parallel counterpart is time consuming and can be done only with the specific knowledge of hardware accelerator architecture, parallel programming or programming environment. The automation of this process makes porting the code much…
Tal Kadosh, Niranjan Hasabnis, Prema Soundararajan, Vy A. Vo + 4 more
Compilation Authors: ['Tal Kadosh' 'Niranjan Hasabnis' 'Prema Soundararajan' 'Vy A. Vo' 'Mihai Capotă' 'Nesreen K. Ahmed' 'Yuval Pinter' 'Gal Oren'] Manual parallelization of code remains a significant challenge due to the complexities of modern software systems and the widespread adoption of multi-core architectures.…
Aleksandr S. Filipchenko
computer system based on Amdahl's law Authors: ['Aleksandr S. Filipchenko'] Abstract The modification of Amdahl's law for the case of increment of processor elements in a computer system is considered. The coefficient k linking accelerations of parallel and parallel specialized computer systems is determined. The…
Andrew Osterhout, Ganesh Gopalakrishnan
It is often difficult to write code that you can ensure will be executed in the right order when programing for parallel compute tasks. Due to the way that today's parallel compute hardware, primarily Graphical Processing Units (GPUs), allows you to write code that needs to be executed on more threads than the device…
Patrick Mukala
| Article Info | ABSTRACT | | --- | --- | | | A myriad of applications ranging from engineering and scientific | | | simulations, image and signal processing as well as high-sensitive data | | | retrieval demand high processing power reaching up to teraflops for their | | | efficient execution. While a standard serial…