22 papers · ranked by Valyu relevance
Harisankar Sadasivan, Milos Maric, Eric Dawson, Vishanth Iyer + 2 more
Long read sequencing technology is becoming increasingly popular for Precision Medicine applications like variant calling from Whole Genome Sequencing (WGS) and for metagenomics applications like microbial abundance estimation. Minimap2 is the state-of-the-art aligner and mapper used by the leading long read sequencing…
Muhammad Osama, Serban D. Porumbescu, John D. Owens
We propose a GPU fine-grained load-balancing abstraction that decouples load balancing from work processing and aims to support both static and dynamic schedules with a programmable interface to implement new load-balancing schedules. Prior to our work, the only way to unleash the GPU's potential on irregular problems…
Hsu-Tzu Ting, Jerry Chou, Ming-Hung Chen, I-Hsin Chung
—Modern GPU workloads increasingly demand efficient resource sharing, as many jobs do not require the full capacity of a GPU. Among sharing techniques, NVIDIA's Multi-Instance GPU (MIG) offers strong resource isolation by enabling hardware-level GPU partitioning. However, leveraging MIG effectively introduces new…
Youhe Jiang, Haoxu Wang, Haotong Bao, Kai Jiang + 4 more
Streaming video generation is emerging as a new serving workload in which users interact with long-lived sessions that generate video progressively, chunk by chunk. Unlike offline video generation or typical LLM serving, streaming video generation must preserve session state across active and idle periods, repeatedly…
Maksudul Alam, Kalyan Perumalla
Synthetically generated, large graph networks serve as useful proxies to real-world networks for many graph-based applications. The ability to generate such networks helps overcome several limitations of real-world networks regarding their number, availability, and access. Here, we present the design, implementation…
Juechu Dong, Xueshen Liu, Harisankar Sadasivan, Sriranjani Sitaraman + 1 more
Long-read DNA sequencing is becoming increasingly popular for genetic diagnostics. Minimap2 is the state-of-the-art long-read aligner. However, Minimap2’s chaining step is slow on the CPU and takes 40-68% of the time especially for long DNA reads. Prior works in accelerating Minimap2 either lose mapping accuracy, are…
Chou-Ying Hsieh, Po-Chieh Lin, Sy‐Yen Kuo
Graphs on GPUs Authors: ['Chou-Ying Hsieh' 'Po-Chieh Lin' 'Sy‐Yen Kuo'] Abstract. The push-relabel algorithm is an efficient algorithm that solves the maximum flow/ minimum cut problems of its affinity to parallelization. As the size of graphs grows exponentially, researchers have used Graphics Processing Units (GPUs)…
Hao Chen Gui, Lin Hu, Rui Chen, Mingxiao Huang + 3 more
Tiling Authors: ['Hao Chen Gui' 'Lin Hu' 'Rui Chen' 'Mingxiao Huang' 'Yan Yin' 'Jin Yang' 'Yong Wu'] 3D Gaussian Splatting (3DGS) is increasingly attracting attention in both academia and industry owing to its superior visual quality and rendering speed. However, training a 3DGS model remains a time-intensive task…
Paul Cardosi, Bérenger Bramas, Bilal Alatas
Parallelization is needed everywhere, from laptops and mobile phones to supercomputers. Among parallel programming models, task-based programming has demonstrated a powerful potential and is widely used in high-performance scientific computing. Not only does it allow efficient parallelization across distributed…
Abedalmuhdi Almomany, Muhammed Sutcu, Babul Salam K. S. M. Kader Ibrahim, Alexandre Bonatto
'Babul Salam K. S. M. Kader Ibrahim' 'Alexandre Bonatto'] Particle-in-cell (PIC) simulation serves as a widely employed method for investigating plasma, a prevalent state of matter in the universe. This simulation approach is instrumental in exploring characteristics such as particle acceleration by turbulence and…
Sarita Simaiya, Umesh Kumar Lilhore, Yogesh Kumar Sharma, K. B. V. Brahma Rao + 4 more
'K. B. V. Brahma Rao' 'V. V. R. Maheswara Rao' 'Anupam Baliyan' 'Anchit Bijalwan' 'Roobaea Alroobaea'] Virtual machine (VM) integration methods have effectively proven an optimized load balancing in cloud data centers. The main challenge with VM integration methods is the trade-off among cost effectiveness, quality of…
Quim Aguado-Puig, Max Doblas, Christos Matzoros, Antonio Espinosa + 3 more
Advances in genomics and sequencing technologies demand faster and more scalable analysis methods that can process longer sequences with higher accuracy. However, classical pairwise alignment methods, based on dynamic programming (DP), impose impractical computational requirements to align long and noisy sequences like…
Altaf Hussain, Muhammad Aleem, Atiq Ur Rehman, Umer Arshad + 1 more
'Michele Pasqua'] Cloud computing provides an opportunity to gain access to the large-scale and high-speed resources without establishing your own computing infrastructure for executing the high-performance computing (HPC) applications. Cloud has the computing resources (i.e., computation power, storage, operating…
Varun C. M, Anto Kumar R. P, Paulraj D, Priyanka P. S + 1 more
To increase cloud computing utilization and performance, efficient load balancing and resource distribution techniques are essential. Dynamic load balancing and resource allocation in cloud systems is necessary due to a number of reasons, but this is not an easy and straightforward task. The primary goal of dynamic…
Yousef Sanjalawe, Salam Fraihat, Salam Al-E’mari, Mosleh Abualhaj + 3 more
'Sharif Makhadmeh' 'Emran Alzubi' 'Davide La Torre'] The increasing dependence on cloud computing as a cornerstone of modern technological infrastructures has introduced significant challenges in resource management. Traditional load-balancing techniques often prove inadequate in addressing cloud environments’ dynamic…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated quadratic…
Authors not listed
Modeling multimetallic systems efficiently enables faster prediction of desirable chemical properties and design of new materials. This work describes an initial implementation for performing multireference wave function method localized active space self-consistent field (LASSCF) calculations through the use of…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Authors not listed
The complete active space self-consistent field (CASSCF) method is essential for describing complex photochemical processes, but its application in ab initio molecular dynamics is often limited by the computational cost associated with four-center two-electron repulsion integrals (ERIs). We present the first…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…
Madushanka Manathunga, Hasan Metin Aktulga, Andreas W. Goetz, Kenneth M. Merz + 1 more
We have ported and optimized the GPU accelerated QUICK and AMBER based ab initio QM/MM implementation on AMD GPUs. This encompasses the entire Fock matrix build and force calculation in QUICK including one-electron integrals, two-electron repulsion integrals, exchange-correlation quadrature, and linear algebra…
Páll Melsted, Elís Mar Guðnýjarson, Jóhannes Nordal
We present a GPU implementation of kallisto for RNA-seq transcript quantification. By redesigning the core algorithms: pseudoalignment, equivalence class intersection, and the EM algorithm; for massively parallel execution on GPUs, we achieve a 30–50× speedup over multithreaded CPU kallisto. On a benchmark of 100…