13 papers · ranked by Valyu relevance
Jovan Blanuša, Paolo Ienne, Kubilay Atasu
Enumerating simple cycles has important applications in computational biology, network science, and financial crime analysis. In this work, we focus on parallelising the state-of-the-art simple cycle enumeration algorithms by Johnson and Read-Tarjan along with their applications to temporal graphs. To our knowledge, we…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
H.K. Al-Mahdawi, A. I Sidikova, Hussein Alkattan, Mostafa Abotaleb + 2 more
'Ammar Kadi' 'El-Sayed M El-kenawy'] We considered in this work the linear operator equation and used the Landweber iterative method as an iterative solver. After that, we used the multigrid method as an optimization method for obtaining an approximation solution with a highly accurate and fast process. A new parallel…
Mohammad Javad Khani, Mahmood Ahmadi
Cyclic Redundancy Check (CRC) remains one of the most widely used error-detection mechanisms in communication, storage, and embedded systems. However, conventional software CRC implementations suffer from inherent sequential dependencies that limit efficient utilization of modern multi-core processors. This paper…
Yang Yu
Sequence Authors: ['Yang Yu'] Recent advances in reasoning models have demonstrated significant improvements in accuracy, particularly for complex tasks such as mathematical reasoning, by employing detailed and comprehensive reasoning processes. However, generating these lengthy reasoning sequences is computationally…
Peiyu Zong, Wenpeng Deng, Jian Liu, Jue Ruan
Farrar recommends employing a stripe method to enhance the implementation of in-sequence parallelism, resulting in significant improvements in parallel performance through the utilization of the SIMD instruction set. This method involves reorganizing the originally sequential computation. The length of each stripe…
Donald S. Ene, V.I.E Anireh
- Evaluating how well a whole system or set of subsystems performs is one of the primary objectives of performance testing. We can tell via performance assessment if the architecture implementation meets the design objectives. Performance evaluations of several parallel algorithms are compared in this study. Both…
Henrik Valter, Axel Karlsson, Miquel Pericàs
OpenMP is the de facto API for parallel programming in HPC applications. These programs are often computed in data centers, where energy consumption is a major issue. Whereas previous work has focused almost entirely on performance, we here analyse aspects of OpenMP from an energy consumption perspective. This analysis…
Mohammed Abutaha, Islam Amar, Salman AlQahtani, Xiaowei Li + 2 more
Encrypting pictures quickly and securely is required to secure image transmission over the internet and local networks. This may be accomplished by employing a chaotic scheme with ideal properties such as unpredictability and non-periodicity. However, practically every modern-day system is a real-time system, for which…
Temitayo Adefemi
—Parallelization has become a cornerstone of modern computing, influencing everything from high-performance supercomputers to everyday mobile devices. This paper presents a comprehensive guide on the fundamentals of parallelization that every computer scientist should know, beginning with a historical perspective that…
Alvaro Estebanez, Diego R. Llanos, David Orden, Belen Palop + 1 more
'Rafael Sachetto Oliveira'] Loops are a rich source of parallelism. Unfortunately, many loops cannot be safely parallelized at compile time because the compiler is not able to guarantee that there will be no dependence violations. Thread-Level Speculation (TLS) techniques, either hardware or software-based, allow the…
Krzysztof Stuglik, Piotr Listkiewicz, Mateusz Kulczyk, Marcin Pietroń
'Marcin Pietroń'] Manual translation of the algorithms from sequential version to its parallel counterpart is time consuming and can be done only with the specific knowledge of hardware accelerator architecture, parallel programming or programming environment. The automation of this process makes porting the code much…
Sirilak Ketchaya, Apisit Rattanatranurak
Quicksort is an important algorithm that uses the divide and conquer concept, and it can be run to solve any problem. The performance of the algorithm can be improved by implementing this algorithm in parallel. In this paper, the parallel sorting algorithm named the Multi-Deque Partition Dual-Deque Merge Sorting…