15 papers · ranked by Valyu relevance
Alvaro Estebanez, Diego R. Llanos, David Orden, Belen Palop + 1 more
'Rafael Sachetto Oliveira'] Loops are a rich source of parallelism. Unfortunately, many loops cannot be safely parallelized at compile time because the compiler is not able to guarantee that there will be no dependence violations. Thread-Level Speculation (TLS) techniques, either hardware or software-based, allow the…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
Yao Xiao, Shahin Nazarian, Paul Bogdan
The urgent need for low latency, high-compute and low power on-board intelligence in autonomous systems, cyber-physical systems, robotics, edge computing, evolvable computing, and complex data science calls for determining the optimal amount and type of specialized hardware together with reconfigurability capabilities.…
Clément Flint, Ludovic Paillat, Bérenger Bramas, Muhammad Aleem
High-performance computing (HPC) relies increasingly on heterogeneous hardware and especially on the combination of central and graphical processing units. The task-based method has demonstrated promising potential for parallelizing applications on such computing nodes. With this approach, the scheduling strategy…
Xiaodong Yu, Viktor Nikitin, Daniel J. Ching, Selin Aslan + 2 more
'Doğa Gürsoy' 'Tekin Biçer'] While the advances in synchrotron light sources, together with the development of focusing optics and detectors, allow nanoscale ptychographic imaging of materials and biological specimens, the corresponding experiments can yield terabyte-scale volumes of data that can impose a heavy burden…
Emad Alamoudi, Felipe Reck, Nils Bundgaard, Frederik Graw + 4 more
'Lutz Brusch' 'Jan Hasenauer' 'Yannik Schälte' 'Abel C.H. Chen'] Approximate Bayesian Computation (ABC) is a widely applicable and popular approach to estimating unknown parameters of mechanistic models. As ABC analyses are computationally expensive, parallelization on high-performance infrastructure is often…
Mateusz Gruzewski, Marek Palkowski, Ramon Antonio Rodriges Zalipynis
In this article, we present an efficient and concise OpenMP implementation of the Nussinov RNA folding algorithm, a well-known representative of non-serial polyadic dynamic programming (NPDP). Our goal is to develop an optimized implementation that can serve as a template for related dynamic programming applications.…
Paul Cardosi, Bérenger Bramas, Bilal Alatas
Parallelization is needed everywhere, from laptops and mobile phones to supercomputers. Among parallel programming models, task-based programming has demonstrated a powerful potential and is widely used in high-performance scientific computing. Not only does it allow efficient parallelization across distributed…
Vladislav Skorpil, Vaclav Oujezsky, Arcangelo Castiglione, Gianni D’Angelo
'Gianni D’Angelo'] This paper presents an implementation of the parallelization of genetic algorithms. Three models of parallelized genetic algorithms are presented, namely the Master-Slave genetic algorithm, the Coarse-Grained genetic algorithm, and the Fine-Grained genetic algorithm. Furthermore, these models are…
Yao Xiao, Guixiang Ma, Nesreen K. Ahmed, Mihai Capotă + 3 more
'Theodore L. Willke' 'Shahin Nazarian' 'Paul Bogdan'] Recent technological advances have contributed to the rapid increase in algorithmic complexity of applications, ranging from signal processing to autonomous systems. To control this complexity and endow heterogeneous computing systems with autonomous programming and…
Sikao Guo, Nenad Korolija, Kent Milfeld, Adip Jhaveri + 3 more
Particle-based reaction-diffusion models offer a high-resolution alternative to the continuum reaction-diffusion approach, capturing the discrete and volume-excluding nature of molecules undergoing stochastic dynamics. These methods are thus uniquely capable of simulating explicit self-assembly of particles into…
Hammad Ather, Sophie Berkman, Giuseppe Cerati, Matti J. Kortelainen + 9 more
Traditionally, high energy physics (HEP) experiments have relied on x86 CPUs for the majority of their significant computing needs. As the field looks ahead to the next generation of experiments such as DUNE and the High-Luminosity LHC, the computing demands are expected to increase dramatically. To cope with this…
Dipanwita Guhathakurta, Fatemeh Rastgar, M. Aditya Sharma, K. Madhava Krishna + 1 more
'K. Madhava Krishna' 'Arun Kumar Singh'] We present a joint multi-robot trajectory optimizer that can compute trajectories for tens of robots in aerial swarms within a small fraction of a second. The computational efficiency of our approach is built on breaking the per-iteration computation of the joint optimization…
Xinyu Qi, Yuchen Yang, Linlin Tian, Zhenming Wang + 1 more
The combination of Cartesian grid and the adaptive mesh refinement (AMR) technology is an effective way to handle complex geometry and solve complex flow problems. Some high-efficiency Cartesian-based AMR libraries have been developed to handle dynamic changes of the grid in parallel but still can not meet the unique…
Cristian Vidal-Silva, Vannessa Duarte, Jesennia Cárdenas-Cobo, Iván Veas
Parallel computing is a current algorithmic approach to looking for efficient solutions; that is, to define a set of processes in charge of performing at the same time the same task. Advances in hardware permit the massification of accessibility to and applications of parallel computing. Nonetheless, some algorithms…