21 papers · ranked by Valyu relevance
Thomas Huber, Swaroop Pophale, Nolan Baker, M. H. Carr + 8 more
'Jaydon Reap' 'Kristina Holsapple' 'Joshua Hoke Davis' 'T. Burnus' 'Seyong Lee' 'David E. Bernholdt' 'Sunita Chandrasekaran'] Abstract—The OpenMP language continues to evolve with every new specification release, as does the need to validate and verify the new features that have been implemented by the different…
Ke Du, Anshu Sharma, Liyi Li, William Mansky
OpenMP is a popular parallelization framework that lets users transform sequential code into parallel code with a few simple annotations. Unfortunately, it is also easy to inadvertently introduce errors by adding OpenMP pragmas into otherwise correct programs, including both logic errors and race conditions. We present…
Nizar Alhafez, Ahmad Kurdi
—This paper presents a comprehensive comparison of three dominant parallel programming models in High Performance Computing (HPC): Message Passing Interface (MPI), Open Multi-Processing (OpenMP), and Compute Unified Device Architecture (CUDA). As computational demands grow exponentially across scientific and industrial…
Gaogao Liu, Wenbo Yang, Peng Li, Guodong Qin + 6 more
'Youming Wang' 'Shuai Wang' 'Ning Yue' 'Dongjie Huang' 'Luís Castedo Ribas'] The data volume and computation task of MIMO radar is huge; a very high-speed computation is necessary for its real-time processing. In this paper, we mainly study the time division MIMO radar signal processing flow, propose an improved MIMO…
Hervé Yviquel, Márcio Pereira, Emílio Francesquini, Guilherme Valarini + 8 more
'Guilherme Valarini' 'Gustavo Leite' 'Pedro Rosso' 'Rodrigo Ceccato' 'Carla Cusihualpa' 'Vitoria Dias' 'Sandro Rigo' 'Alan A. V. B. Souza' 'Guido Araújo'] Despite the various research initiatives and proposed programming models, efficient solutions for parallel programming in HPC clusters still rely on a complex…
Johannes Blühdorn, Max Sagebaum, Nicolas R. Gauger
We present the new software OpDiLib, a universal add-on for classical operator overloading AD tools that enables the automatic differentiation (AD) of OpenMP parallelized code. With it, we establish support for OpenMP features in a reverse mode operator overloading AD tool to an extent that was previously only reported…
Nhan Ly-Trong, Giuseppe M.J. Barca, Bui Quang Minh
Sequence simulation plays a vital role in phylogenetics with many applications, such as evaluating phylogenetic methods, testing hypotheses, and generating training data for machine-learning applications. We recently introduced a new simulator for multiple sequence alignments called AliSim, which outperformed existing…
Petros Voudouris, Per Stenström, Risat Pathan
Heterogeneous multiprocessors can offer high performance at low energy expenditures. However, to be able to use them in hard real-time systems, timing guarantees need to be provided, and the main challenge is to determine the worst-case schedule length (also known as makespan) of an application. Previous works that…
Hui Zhou, Ken Raffenetti, Junchao Zhang, Yanfei Guo + 1 more
MPI+Threads, embodied by the MPI/OpenMP hybrid programming model, is a parallel programming paradigm where threads are used for on-node shared-memory parallelization and MPI is used for multi-node distributed-memory parallelization. OpenMP provides an incremental approach to parallelize code, while MPI, with its…
John Kruper, Ariel Rokem
Tractography based on diffusion-weighted MRI (dMRI) is the predominant in vivo method for mapping the brain’s white matter. However, it is also one of the most computationally demanding steps in neuroimaging data analysis-requiring the generation and filtering of millions of streamlines per subject. Over the past…
Nhan Ly-Trong, Giuseppe M J Barca, Bui Quang Minh, Russell Schwartz
This paper introduces AliSim-HPC, a high-performance sequence simulator for phylogenetics. We present two multi-threading algorithms to simulate a single large gap-free alignment with OpenMP, and an embarrassingly parallel scheme to simulate many alignments (with/without gaps) with MPI on a distributed-memory system.…
Wilfried Agbeto, Camille Coti, Vladimir Reinharz
Advances in graph algorithmics have allowed in-depth study of many natural objects from molecular biology or chemistry to social networks. Particularly in molecular biology and cheminformatics, understanding complex structures by identifying conserved sub-structures is a key milestone towards the artificial design of…
Paul Cardosi, Bérenger Bramas, Bilal Alatas
Parallelization is needed everywhere, from laptops and mobile phones to supercomputers. Among parallel programming models, task-based programming has demonstrated a powerful potential and is widely used in high-performance scientific computing. Not only does it allow efficient parallelization across distributed…
Bence Ladóczki, László Gyevi-Nagy, Péter R. Nagy, Mihály Kállay
and Large-Scale Explicitly Correlated CCSD(T) Computations via a Reduced-Cost and Parallel Implementation Authors: ['Bence Ladóczki' 'László Gyevi-Nagy' 'Péter R. Nagy' 'Mihály Kállay'] Parallel algorithms to accelerate explicitly correlated second-order Mo̷ller-Plesset (MP2) and coupled-cluster singles and doubles…
Lu Chen, Dejun Teng, Tian Zhu, Jun Kong + 4 more
The human body is made up of about 37 trillion cells (adults). Each cell has its own unique role and is affected by its neighboring cells and environment. The NIH Human BioMolecular Atlas Program (HuBMAP) aims at developing a 3D atlas of human body consisting of organs, vessels, tissues to singe cells with all 3D…
Krzysztof M Ocetkiewicz, Cezary Czaplewski, Henryk Krawczyk, Agnieszka G Lipska + 5 more
Our parallelized CUDA code led to significant speed-ups over the CPU OpenMP version. In [btad391-F1], we visualized the average (out of 3) execution times for input datasets of various sizes measured on a modern workstation with 2× AMD EPYC 7313 CPUs @3.0 GHz (2 × 16 physical cores), 8× NVIDIA A100 40GB GPU, and 4TB…
Authors not listed
The era of exascale computing presents both exciting opportunities and unique challenges for quantum mechanical simulations. While the transition from petaflops to exascale computing has been marked by a steady increase in computational power, the shift towards heterogeneous architectures, particularly the dominant…
Willem A.M. Wybo, Leander Ewert, Charl Linssen, Pooja Babu + 3 more
While the implementation of learning and memory in the brain is governed in large part by subcellular mechanims in the dendrites of neurons, large-scale network simulations featuring such processes remain challenging to achieve. This can be attributed to a lack of appropriate software tools, as neuroscientific…
Sikao Guo, Nenad Korolija, Kent Milfeld, Adip Jhaveri + 3 more
Particle-based reaction-diffusion models offer a high-resolution alternative to the continuum reaction-diffusion approach, capturing the discrete and volume-excluding nature of molecules undergoing stochastic dynamics. These methods are thus uniquely capable of simulating explicit self-assembly of particles into…
David S. Cerutti, Rafal Wiewiora, Simon Boothroyd, Woody Sherman
The Structure and TOpology Replica Molecular Mechanics (STORMM) code is a next-generation molecular simulation engine and associated libraries optimized for performance on fast, multicore central processor units (CPUs) and graphics processing units (GPUs) with independent memory and tens of thousands of threads. STORMM…
Authors not listed
The increasing importance and predictive power of modern molecular modeling, driven by physics- and machine learning-based methods, necessitates a new collaborative architecture to replace the isolated, traditional model of software development. The traditional approach often led to redundant engineering effort, high…