13 papers · ranked by Valyu relevance
Haymo Kutschbach
This work introduces a self-optimizing virtual processor (VP) for numerical array programs that shifts parallelization from a manual developer task to a cooperative, agent-like runtime mechanism. Instead of relying on centralized task-graph scheduling, static compiler optimization, or explicitly annotated parallel…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Stavrinides, Georgios L., Karatza, Helen D.
With the explosive growth of big data, workloads tend to get more complex and computationally demanding. Such applications are processed on distributed interconnected resources that are becoming larger in scale and computational capacity. Data-intensive applications may have different degrees of parallelism and must…
Jinyang Wang, Zhugang Wang, Di Liu, Isabelle Ledoux-Rak
Ultra-high sampling rates in coherent optical front-ends increasingly exceed the processing capabilities of real-time baseband processors, creating a bottleneck in coherent free-space optical communication systems. We propose a unified state-space framework to systematically parallelize digital signal processing (DSP)…
Tasmia Jannat, Michael Gowanlock, Satish Puri
The growing volume of data in scientific domains has made spatial query processing increasingly challenging due to high data transfer costs across the memory hierarchy and limited memory bandwidth. To address these bottlenecks and reduce the energy consumed on data movement, this work explores Processing-in-Memory…
Ran Ginosar
I have greatly enjoyed spending many years in studying parallel computing. My journey goes thorough MP-C, PLURAL, Async Plural, HAL, RC64 and more. As a PhD student at Princeton I studied a combination of shared memory and message passing, motivated by algorithms and the ease of programming. While at the Technion, a…
Kenshin Obi, Ryo Yoshinaka, Fujimoto, Hiroshi + 1 more
—In recent years, the complexity and scale of embedded systems, especially in the rapidly developing field of autonomous driving systems, have increased significantly. This has led to the adoption of software and hardware approaches such as Robot Operating System (ROS) 2 and multi-core processors. Traditional manual…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
Drew E. Winters
Studying flexible, adaptive transitions between cognitive tasks and serial-parallel processing under changing task demands has been a central focus for understanding human cognition. Advances in neuroimaging analysis have improved the ability to link cognition with brain function, providing a foundation for developing…
Farzad Razi, Mehran Moghadam, Sercan Aygun, M. Hassan Najafi + 1 more
Today's high-performance architectures are increasingly constrained by data movement latency and energy overhead, as the slowdown of single-core performance scaling coincides with the rise of highly data-intensive workloads. In-memory architectures have emerged as a complementary solution to conventional von Neumann…
Shujun Peng, Xinhan Lin, Yu Zhang, Yuheng Xiao + 1 more
Parallel scan is a fundamental primitive widely used in a broad range of workloads, including parallel sorting, graph algorithms, and sampling in large language model inference. Although GPU-optimized parallel scan algorithms have been extensively studied, their reliance on vector units makes them inefficient on modern…
Rubén Langarita, Jesús Alastruey-Benedé, Pablo Ibáñez, Santiago Marco‐Sola + 2 more
—Multiple HPC applications are often bottlenecked by compute-intensive kernels implementing complex dependency patterns (data-dependency bound). Traditional general-purpose accelerators struggle to effectively exploit fine-grain parallelism due to limitations in implementing convoluted data-dependency patterns (like…
Drew E. Winters
Studying flexible, adaptive transitions between cognitive tasks and serial-parallel processing under changing task demands has been central to understanding human cognition. Advances in neuroimaging analysis have improved the ability to link cognition with brain function, motivating methods that characterize dynamic…