19 papers · ranked by Valyu relevance
Mamirov, Akhmadillo
—GPU clusters have become essential for training and deploying modern AI systems, yet real deployments continue to report average utilization near 50%. This inefficiency is largely caused by fragmentation, heterogeneous workloads, and the limitations of static scheduling policies. This work presents a systematic…
Seokjin Go, Marko Scrbak, Ephrem Wu, Srilatha Manne + 1 more
In distributed Mixture-of-Experts (MoE) inference, input-dependent token routing interacts with GPU performance variability to create persistent stragglers under synchronized execution, where the slowest GPU determines layer latency. This performance variability is inherent to modern accelerators: manufacturing…
Youhe Jiang, Haoxu Wang, Haotong Bao, Kai Jiang + 4 more
Streaming video generation is emerging as a new serving workload in which users interact with long-lived sessions that generate video progressively, chunk by chunk. Unlike offline video generation or typical LLM serving, streaming video generation must preserve session state across active and idle periods, repeatedly…
Shruti Dongare, Redwan Ibne Seraj Khan, Hadeel Albahar, Nannan Zhao + 2 more
Modern cloud platforms increasingly host large-scale deep learning (DL) workloads, demanding high-throughput, low-latency GPU scheduling. However, the growing heterogeneity of GPU clusters and limited visibility into application characteristics pose major challenges for existing schedulers, which often rely on offline…
Kevin Garner, Chander Sadasivan, Nikos Chrisochoides
This paper presents two performance optimization techniques for a mesh adaptation method that is designed to help streamline the discretization of complex vascular geometries within the numerical modeling process. This method is integrated into a pipeline with an image-to-mesh conversion tool to generate adaptive…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Authors not listed
Recent advances in machine learning force fields (MLFF) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference datasets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and…
Varun C. M, Anto Kumar R. P, Paulraj D, Priyanka P. S + 1 more
To increase cloud computing utilization and performance, efficient load balancing and resource distribution techniques are essential. Dynamic load balancing and resource allocation in cloud systems is necessary due to a number of reasons, but this is not an easy and straightforward task. The primary goal of dynamic…
Nomula Divya, M. N. V. Kiranbabu, G. Charles Babu
With the exponential growth of multi-cloud, spearheaded by latency-sensitive workloads, edge inclusion and heterogeneous resource pools; there is a dire need for scheduling methods which are energy neutral as well as robust against dynamic operational stress. In this work, a new bi-directional optimization-prediction…
Authors not listed
The complete active space self-consistent field (CASSCF) method is essential for describing complex photochemical processes, but its application in ab initio molecular dynamics is often limited by the computational cost associated with four-center two-electron repulsion integrals (ERIs). We present the first…
Amin Saberi, Bin Wan, Kevin J. Wischnewski, Kyesam Jung + 6 more
Brain network modeling uses computer simulations to infer about latent neural properties at micro- and mesoscales by fitting brain dynamic models to empirical data of individual subjects or groups. However, computational costs of (individualized) model fitting is a major bottleneck, limiting the practical feasibility…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…
Lin Cai, Xuzhong Qu, Hang Zhou, Ning Li + 8 more
High-resolution large-volume biological imaging techniques are now widely used in biological research, but inefficiencies in data processing and visualization persist due to bottlenecks in loading/saving pipelines and limitations of conventional pyramid formats. To address these issues, we have decoupled data loading…
Aryan Kumar, Punit Gupta, Rohit Verma, Asaad Ahmed Gad Elrab Ahmed
Fog computing minimizes latency and bandwidth consumption by processing data near the source, but it has the challenge of workload balancing across the dynamic and resource limited fog nodes. Uneven task assignment can result in bottlenecks, idle resources, and lowered Quality of Service (QoS) standards. In this work…
Matthew Leach, Peter Heywood, Alexander G. Fletcher, Paul Richmond
Chaste is an open-source C++ library providing a general-purpose framework for cell-based simulations of biological tissues. It has been applied to a wider range of biological processes, including morphogenesis, carcinogenesis, and wound healing. Such simulations often involve numerous mechanical interactions between…
Authors not listed
Machine Learning Interatomic Potentials (MLIPs), trained with Quantum Mechanics data, can model potential energy surfaces for molecular systems with very high accuracy and extreme speedups compared to reference quantum calculations, offering a powerful tool for studying complex chemical and biological systems. This…
Michail Patsakis, Alexandros Tzanakakis, Ilias Georgakopoulos-Soares
Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2’s weights and attention cache to four…
Emre Green, Adil Mardinoglu
Whole-genome sequencing (WGS) has transformed clinical diagnostics, yet variant annotation remains a computational bottleneck. The Variant Effect Predictor (VEP) integrates pathogenicity predictors and population databases essential for ACMG/AMP variant classification, but these annotation plugins are fundamentally…
Páll Melsted, Elís Mar Guðnýjarson, Jóhannes Nordal
We present a GPU implementation of kallisto for RNA-seq transcript quantification. By redesigning the core algorithms: pseudoalignment, equivalence class intersection, and the EM algorithm; for massively parallel execution on GPUs, we achieve a 30–50× speedup over multithreaded CPU kallisto. On a benchmark of 100…