23 papers · ranked by Valyu relevance
Emanuele Del Sozzo, Martin Fleming, Kenneth Flamm, Neil Thompson
—Graphics Processing Units (GPUs) are the state-ofthe-art architecture for essential tasks, ranging from rendering 2D/3D graphics to accelerating workloads in supercomputing centers and, of course, Artificial Intelligence (AI). As GPUs continue improving to satisfy ever-increasing performance demands, analyzing past…
Alessio Suriano, Stefano Truzzi, Agnese Costa, Marco Rossazza + 4 more
a UNITO - Università degli Studi di Torino, Dipartimento di Fisica, via Giuria, 1, Torino, 10100, Italy b INAF - Istituto Nazionale di Astrofisica, Osservatorio Astrofisico di Torino, Strada Osservatorio, 20, Pino Torinese, 10025, Italy c ICSC - Italian Research Center on High Performance Computing, Big Data and…
Yunhao Wang, Qite Wang, Yan Wang, Mikhail Sheremet
This paper presents an efficient and high-order WENO-based Upwind Rotated Lattice Boltzmann Flux Solver (WENO-URLBFS) on graphics processing units (GPUs) for simulating three-dimensional (3D) compressible flow problems. The proposed approach extends the baseline Rotated Lattice Boltzmann Flux Solver (RLBFS) by…
Antonina Dobrowolska, Julian Świerczyński, Paweł Tecmer, Emil Sujkowski + 5 more
Python Frameworks for Next-Generation GPUs: A Comparative Study of CuPy and PyTorch on the Hopper and Grace Hopper Architecture Authors: Antonina Dobrowolska, Julian Świerczyński, Paweł Tecmer, Emil Sujkowski, Somayeh Ahmadkhani, Grzegorz Mazur, Klemens Noga, Jeff Hammond, Katharina Boguslawski In this work, we…
Hsu-Tzu Ting, Jerry Chou, Ming-Hung Chen, I-Hsin Chung
—Modern GPU workloads increasingly demand efficient resource sharing, as many jobs do not require the full capacity of a GPU. Among sharing techniques, NVIDIA's Multi-Instance GPU (MIG) offers strong resource isolation by enabling hardware-level GPU partitioning. However, leveraging MIG effectively introduces new…
Matthew Leach, Peter Heywood, Alexander G. Fletcher, Paul Richmond
Chaste is an open-source C++ library providing a general-purpose framework for cell-based simulations of biological tissues. It has been applied to a wider range of biological processes, including morphogenesis, carcinogenesis, and wound healing. Such simulations often involve numerous mechanical interactions between…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Yuang Yan, Ian Karlin, Ryan Grant
For NVIDIA GPUs, CUDA is the primary interface through which applications orchestrate GPU execution, yet much of the logic that realizes CUDA operations resides in NVIDIA's closed-source userspace driver. As a result, the translation from high-level CUDA APIs to low-level hardware commands remains opaque, limiting both…
Páll Melsted, Elís Mar Guðnýjarson, Jóhannes Nordal
We present a GPU implementation of kallisto for RNA-seq transcript quantification. By redesigning the core algorithms: pseudoalignment, equivalence class intersection, and the EM algorithm; for massively parallel execution on GPUs, we achieve a 30–50× speedup over multithreaded CPU kallisto. On a benchmark of 100…
Ayesha Afzal, Kahler, Anna, Georg Hager + 1 more
Molecular dynamics simulations are essential tools in computational biophysics, but their performance depend heavily on hardware choices and configuration. In this work, we presents a comprehensive performance analysis of four NVIDIA GPU accelerators – A40, A100, L4, and L40 – using six representative GROMACS…
S. Bnà, Giuseppe Giaquinto, Ettore Fadiga, Tommaso Zanelli + 1 more
High Performance Computing (HPC) on hybrid clusters represents a significant opportunity for Computational Fluid Dynamics (CFD), especially when modern accelerators are utilized effectively. However, despite the widespread adoption of GPUs, programmability remains a challenge, particularly in open-source contexts. In…
Emre Green, Adil Mardinoglu
Whole-genome sequencing (WGS) has transformed clinical diagnostics, yet variant annotation remains a computational bottleneck. The Variant Effect Predictor (VEP) integrates pathogenicity predictors and population databases essential for ACMG/AMP variant classification, but these annotation plugins are fundamentally…
Rakbin Sung, Seongmi Woo, Dongmin Shin, Junil Kim + 2 more
On the zebrafish dataset, FastSCODE achieved up to a 2532-fold speedup when utilizing three NVIDIA RTX 4090 GPUs ([btaf624-F1]). Even greater acceleration was observed for the CeNGEN dataset, where FastSCODE achieved >6000-fold speedup using four 4090 GPUs ([btaf624-F1]). The actual execution time was reduced from 8383…
Qiujiang Liang, Jun Yang
for Performant Large-Scale Ab Initio Calculations Authors: Qiujiang Liang, Jun Yang Computational acceleration of orbital-invariant local correlation methods on graphics processing units (GPUs) has remained largely unexplored due to substantial algorithmic complexities. The runtime efficiency of GPU-implemented local…
Authors not listed
The complete active space self-consistent field (CASSCF) method is essential for describing complex photochemical processes, but its application in ab initio molecular dynamics is often limited by the computational cost associated with four-center two-electron repulsion integrals (ERIs). We present the first…
Authors not listed
Machine Learning Interatomic Potentials (MLIPs), trained with Quantum Mechanics data, can model potential energy surfaces for molecular systems with very high accuracy and extreme speedups compared to reference quantum calculations, offering a powerful tool for studying complex chemical and biological systems. This…
Marco Savioli, Paolo Calligari, Ugo Locatelli, Gianfranco Bocchinfuso
We introduce GROMODEX, a novel tool designed to optimise GROMACS molecular dynamics (MD) simulations using a structured Design of Experiments (DoE) approach. GROMACS, though efficient, requires extensive tuning of parameters to perform optimally on different hardware and molecular systems. Manual tuning is tedious and…
Authors not listed
We present the next generation of AMP, a neural network potential (NNP) with anisotropic message passing designed to study large biomolecular systems at DFT accuracy in the condensed phase using a multiscale approach similar to quantum-mechanics/molecular-mechanics (QM/MM) with electrostatic embedding. We trained AMPv3…
Authors not listed
Recent advances in machine learning force fields (MLFF) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference datasets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and…
Spencer Starr, Yannik Feldner, Patrick Kopper, Marcel Blind + 5 more
With the recent proliferation of heterogeneous, GPU-accelerated supercomputers, high-order computational fluid dynamics (CFD) simulations of complex, turbulent flows are more accessible than ever. To leverage the computing power of these machines, CFD software must adapt. However, complicating the situation is the…
Authors not listed
We describe a collaborative research project spanning the disciplines of quantum hardware, quantum algorithms, conventional computational chemistry, synthetic medicinal chemistry and life sciences. Our project seeks to demonstrate an impact of quantum computing on human health. It is one of several funded by Wellcome…
Authors not listed
Equivariant graph neural networks have shown remarkable success in molecular property prediction, but their performance on novel molecular geometries remains limited without extensive training data. We present a computationally efficient approach to cross-geometry pretraining for molecular systems that improves…
Authors not listed
Computational chemistry has entered a new era where machine learning (ML) models—particularly graph neural networks and machine learning force fields—routinely deliver quantum mechanical accuracy at classical speeds, scaling to millions of atoms and reshaping workflows in drug discovery, catalysis, and materials…