18 papers · ranked by Valyu relevance
Ke Du, Anshu Sharma, Liyi Li, William Mansky
OpenMP is a popular parallelization framework that lets users transform sequential code into parallel code with a few simple annotations. Unfortunately, it is also easy to inadvertently introduce errors by adding OpenMP pragmas into otherwise correct programs, including both logic errors and race conditions. We present…
Kent Milfeld, Bronis R. de Supinski, Lars Koesterke, Jannis Klinkenberg + 4 more
'Jannis Klinkenberg' 'Lechen Yu' 'Joachim Protze' 'Oscar Hernandez' 'Vivek Sarkar'] Incorrect usage of OpenMP constructs may cause different kinds of defects in OpenMP applications. Most of the existing work focuses on concurrency bugs such as data races and deadlocks, since concurrency bugs are difficult to detect and…
Vibha Rajput, Alok Katiyar
—The aim of parallel computing is to increase an application's performance by executing the application on multiple processors. OpenMP is an API that supports multiplatform shared memory programming model and sharedmemory programs are typically executed by multiple threads. The use of multi threading can enhance the…
Hervé Yviquel, Márcio Pereira, Emílio Francesquini, Guilherme Valarini + 8 more
'Guilherme Valarini' 'Gustavo Leite' 'Pedro Rosso' 'Rodrigo Ceccato' 'Carla Cusihualpa' 'Vitoria Dias' 'Sandro Rigo' 'Alan A. V. B. Souza' 'Guido Araújo'] Despite the various research initiatives and proposed programming models, efficient solutions for parallel programming in HPC clusters still rely on a complex…
Albert Saà-Garriga, David Castells-Rufas, Jordi Carrabina
In this paper, we present OMP2MPI a tool that generates automatically MPI source code from OpenMP. With this transformation the original program can be adapted to be able to exploit a larger number of processors by surpassing the limits of the node level on large HPC clusters. The transformation can also be useful to…
Ananya Muddukrishna, Peter A. Jonsson, Mats Brorsson, Vince Grolmusz
Programmers struggle to understand performance of task-based OpenMP programs since profiling tools only report thread-based performance. Performance tuning also requires task-based performance in order to balance per-task memory hierarchy utilization against exposed task parallelism. We provide a cost-effective method…
Yi Ding, Kai Hu, Kai Wu, Zhenlong Zhao + 1 more
OpenMP, a typical shared memory programming paradigm, has been extensively applied in high performance computing community due to the popularity of multicore architectures in recent years. The most significant feature of the OpenMP 3.0 specification is the introduction of the task constructs to express parallelism at a…
Albert Saà-Garriga, David Castells-Rufas, Jordi Carrabina
We present OMP2HMPP, a tool that, in a first step, automatically translates OpenMP code into various possible transformations of HMPP. In a second step OMP2HMPP executes all variants to obtain the performance and power consumption of each transformation. The resulting trade-off can be used to choose the more convenient…
Johannes Blühdorn, Max Sagebaum, Nicolas R. Gauger
We present the new software OpDiLib, a universal add-on for classical operator overloading AD tools that enables the automatic differentiation (AD) of OpenMP parallelized code. With it, we establish support for OpenMP features in a reverse mode operator overloading AD tool to an extent that was previously only reported…
Pierre Carrier, Bill Long, Richard Walsh, Jef Dawson + 4 more
High Performance Computing (HPC) Best Practice offers opportunities to implement lessons learned in areas such as computational chemistry and physics in genomics workflows, specifically Next-Generation Sequencing (NGS) workflows. In this study we will briefly describe how distributed-memory parallelism can be an…
Nhan Ly-Trong, Giuseppe M.J. Barca, Bui Quang Minh
Sequence simulation plays a vital role in phylogenetics with many applications, such as evaluating phylogenetic methods, testing hypotheses, and generating training data for machine-learning applications. We recently introduced a new simulator for multiple sequence alignments called AliSim, which outperformed existing…
Wei Wang, Lifan Xu, John Cavazos, Howie H. Huang + 2 more
'Tobias Preis'] Recent developments in modern computational accelerators like Graphics Processing Units (GPUs) and coprocessors provide great opportunities for making scientific applications run faster than ever before. However, efficient parallelization of scientific code using new programming tools like CUDA requires…
Eric Wright, Mauricio Ferrato, Alex Bryer, Robert Searles + 2 more
Experimental chemical shifts (CS) from solution and solid state magic-angle-spinning nuclear magnetic resonance spectra provide atomic level information for each amino acid within a protein or protein complex. However, structure determination of large complexes and assemblies based on NMR data alone remains challenging…
Ahmadreza Ghaffarizadeh, Samuel H. Friedman, Shannon M Mumenthaler, Paul Macklin
Many multicellular systems problems can only be understood by studying how cells move, grow, divide, interact, and die. Tissue-scale dynamics emerge from systems of many interacting cells as they respond to and influence their microenvironment. The ideal “virtual laboratory” for such multicellular systems simulates…
Viacheslav Bolnykh, Jógvan Magnus Haugaard Olsen, Simone Meloni, Martin P. Bircher + 3 more
We present a highly scalable DFT-based QM/MM implementation developed within MiMiC, a recently introduced multiscale modeling framework that uses a loose-coupling strategy in conjunction with a multiple-program multiple-data (MPMD) approach. The computation of electrostatic QM/MM interactions is parallelized exploiting…
Authors not listed
The era of exascale computing presents both exciting opportunities and unique challenges for quantum mechanical simulations. While the transition from petaflops to exascale computing has been marked by a steady increase in computational power, the shift towards heterogeneous architectures, particularly the dominant…
Authors not listed
The increasing importance and predictive power of modern molecular modeling, driven by physics- and machine learning-based methods, necessitates a new collaborative architecture to replace the isolated, traditional model of software development. The traditional approach often led to redundant engineering effort, high…
Jingcheng Shen, Jie Mei, Marcus Walldén, Fumihiko Ino
FreeSurfer is among the most widely used suites of software for the study of cortical and subcortical brain anatomy. However, analysis using FreeSurfer can be time-consuming and it lacks support for the graphics processing units (GPUs) after the core development team stopped maintaining GPU-accelerated versions due to…