15 papers · ranked by Valyu relevance
Li, Ruihao, John, Lizy K. + 2 more
Memory allocators hide beneath nearly every application stack, yet their performance footprint extends far beyond their code size. Even small inefficiencies in the allocators ripple through caches and the rest of the memory hierarchy, collectively imposing what operators often call a "datacenter tax". At hyperscale…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Aleix Roca, Vicenç Beltran
The convergence of high-performance computing (HPC) and artificial intelligence (AI) is driving the emergence of increasingly complex parallel applications and workloads. These workloads often combine multiple parallel runtimes within the same application or across co-located jobs, creating scheduling demands that…
Kenshin Obi, Ryo Yoshinaka, Fujimoto, Hiroshi + 1 more
—In recent years, the complexity and scale of embedded systems, especially in the rapidly developing field of autonomous driving systems, have increased significantly. This has led to the adoption of software and hardware approaches such as Robot Operating System (ROS) 2 and multi-core processors. Traditional manual…
Arsham Mikaeili Namini, Ali Saberi, Hamed S Najafabadi, Peter Robinson
Runtime benchmarking demonstrated substantial improvements from both architectural redesign and parallelization ([btag334-F1]). Even without parallelism, GEDI 2.0 ran up to $∼$3 $\times$ faster than the legacy implementation, confirming substantial performance gains from improved algorithms and memory access patterns…
Sjoerd Dost
Concurrent logic programming predates miniKanren, but concurrent implementations of miniKanren have remained largely unexplored. In this work we present a parallel implementation of miniKanren in Go, demonstrating its feasibility and potential for performance improvements. Our approach leverages implicit parallelism…
R. Govind, S. Krishna, Sanchari Sil, B. Srivathsan
Gibbons and Korach studied a fundamental problem in 1997: given an observed sequence of reads and writes of a multi-threaded program, does there exist an interleaving which is sequentially consistent? Apart from applications in testing shared memory implementations, a procedure for this problem is employed in Dynamic…
Sangjin Lee, Sunggon Kim, Yongseok Son, Agbotiname Lucky Imoize
We propose ScaleDefrag, a parallel and asynchronous defragmentation tool that reduces defragmentation time by up to 3.8× compared to e4defrag, while improving scalability on multi-core systems. Flash-based solid-state drives (SSDs) have been widely adopted in various large-scale storage systems including cloud and HPC…
Raaghav Ravishankar, Sandeep Kulkarni, Sathya Peri, Gokarna Sharma + 2 more
A domain extension of a definition refers to broadening the scope of a definition so that it applies to a larger set of cases than originally specified. The notion of lock-free and wait-free computation is designed for the domain of tasks that are completed by a single thread (in competition with other threads). The…
Alexander Alsalihi, Robert M. Flight, Hunter N. B. Moseley
The recount3 online resource provides tens of thousands of uniformly processed RNA-seq samples across human and mouse from major sequencing repositories like the Sequence Read Archive. While access to these datasets has traditionally been centered in the R/Bioconductor ecosystem, the growing prominence of Python in…
Teh Noranis Mohd Aris, Ningning Chen, Norwati Mustapha, Maslina Zolkepli + 1 more
To address the inefficiencies in sample utilization and policy instability in asynchronous distributed reinforcement learning, we propose TPDEB-a dual experience replay framework that integrates prioritized sampling and temporal diversity. While recent distributed RL systems have scaled well, they often suffer from…
Anders Pitman, Cathy Yang, Yi Qiao
Next-generation sequencing now produces whole-genome data in hours, but downstream variant calling remains a multi-hour to multi-day bottleneck that excludes genomic analysis from time-critical clinical settings. GPU acceleration offers a natural path forward — variant calling is inherently parallelizable across…
Kai Xu, Diming Zhang, Xuguo Wang, Alessandra Rizzardi
To address the bottlenecks of missing decision-making closed loop, insufficient experience reuse, and decoupled resource scheduling in industrial LLM deployment, this paper proposes LLM-Conductor, a three-layer collaborative architecture that enables monitoring-feedback autonomous decision-making, structured policy…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…
Authors not listed
We present an open source collection of scripts and programs for the setup, management and evaluation of calculations with the Vienna ab-initio simulation package (VASP), called utils4VASP. It contains 20 independent Python scripts and Fortran programs, all with a unified and intuitive handling concept based on command…