23 papers · ranked by Valyu relevance
Maria Clicia S. Castro, Vanessa S. Silva, Maiana O. C. Costa, Helena S. I. L. Silva + 8 more
Several hundred terabytes of single-cell RNA-seq (scRNA-seq) data are available in public repositories. These data refer to various research projects, from microbial population cells to multiple tissues, involving patients with a myriad of diseases and comorbidities. An increase to several Petabytes of these scRNA-seq…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Authors not listed
This paper presents GLAS (Git-based Lab Automated Scheduler or Get Lab Automation Simplified), an open-source, robust, and highly expandable Git-based architecture designed for laboratory automation. GLAS can be deployed in both partially and fully automated experimental science laboratories, enabling the development…
Paul Cardosi, Bérenger Bramas, Bilal Alatas
Parallelization is needed everywhere, from laptops and mobile phones to supercomputers. Among parallel programming models, task-based programming has demonstrated a powerful potential and is widely used in high-performance scientific computing. Not only does it allow efficient parallelization across distributed…
Parinaz Barakhshan, Rudolf Eigenmann
We compare automatically and manually parallelized NAS Benchmarks in order to identify code sections that differ. We discuss opportunities for advancing automatic parallelizers. We find ten patterns that pose challenges for current parallelization technology. We also measure the potential impact of advanced techniques…
Alvaro Estebanez, Diego R. Llanos, David Orden, Belen Palop + 1 more
'Rafael Sachetto Oliveira'] Loops are a rich source of parallelism. Unfortunately, many loops cannot be safely parallelized at compile time because the compiler is not able to guarantee that there will be no dependence violations. Thread-Level Speculation (TLS) techniques, either hardware or software-based, allow the…
Junichiro Makino, Toshikazu Ebisuzaki, Ryutaro Himeno, Yoshihide Hayashizaki
Rapidly increasing amount of short read data generated by NGSs (new-generation sequencers) calls for the development of fast and accurate read alignment programs. The programs based on hash table (BLAST) and Burrows-Wheeler transform (bwa-mem) are used, and the latter is known to give superior performance. We here…
Temitayo Adefemi
—Parallelization has become a cornerstone of modern computing, influencing everything from high-performance supercomputers to everyday mobile devices. This paper presents a comprehensive guide on the fundamentals of parallelization that every computer scientist should know, beginning with a historical perspective that…
Madhav P. Desai
We investigate the utility of augmenting a microprocessor with a single execution pipeline by adding a second copy of the execution pipeline in parallel with the existing one. The resulting dual-hardware-threaded microprocessor has two identical, independent, single-issue in-order execution pipelines (hardware threads)…
Urmila Shrawankar, Mayuri Joshi
— In multi-core systems, various factors like inter-process communication, dependency, resource sharing and scheduling, level of parallelism, synchronization, number of available cores etc. influence the extent of possible High Performance Computing parallelization. These parameters if not managed to the root level…
Denis Los, Igor Petushkov
Cores Authors: ['Denis Los' 'Igor Petushkov'] Abstract—Nowadays, latency-critical, high-performance applications are parallelized even on power-constrained client systems to improve performance. However, an important scenario of fine-grained tasking on simultaneous multithreading CPU cores in such systems has not been…
Sikao Guo, Nenad Korolija, Kent Milfeld, Adip Jhaveri + 3 more
Particle-based reaction-diffusion models offer a high-resolution alternative to the continuum reaction-diffusion approach, capturing the discrete and volume-excluding nature of molecules undergoing stochastic dynamics. These methods are thus uniquely capable of simulating explicit self-assembly of particles into…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Juhyeon Park, Heoncheol Lee, Hyuck-Hoon Kwon, Yeji Hwang + 2 more
'Wonseok Choi' 'Wei Yi'] This paper addresses the problem of tracking a high-speed ballistic target in real time. Particle swarm optimization (PSO) can be a solution to overcome the motion of the ballistic target and the nonlinearity of the measurement model. However, in general, particle swarm optimization requires a…
Kecong Tang, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang
Durbin’s positional Burrows-Wheeler transform (PBWT) enables algorithms with the optimal time complexity of O(MN) for reporting all vs all haplotype matches in a population panel with M haplotypes and N variant sites. However, even this efficiency may still be too slow when the number of haplotypes reaches millions. To…
Cristian Vidal-Silva, Vannessa Duarte, Jesennia Cárdenas-Cobo, Iván Veas
Parallel computing is a current algorithmic approach to looking for efficient solutions; that is, to define a set of processes in charge of performing at the same time the same task. Advances in hardware permit the massification of accessibility to and applications of parallel computing. Nonetheless, some algorithms…
Jesper Larsson Träff
These lecture notes are designed to accompany an imaginary, virtual, undergraduate, one or two semester course on fundamentals of Parallel Computing as well as to serve as background and reference for graduate courses on High-Performance Computing, parallel algorithms and shared-memory multiprocessor programming. They…
Authors not listed
Protein conformational landscapes contain the functionally relevant information useful for understanding biological processes. Mapping out conformational landscapes provides valuable insights into protein behaviors and biological phenomena, and has relevance to therapeutic design. While experimental structural biology…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated quadratic…
Andrew Simmonett, Bernard Brooks, Thomas Darden
Evaluation of noncovalent electrostatic interactions is the dominant bottleneck in classical molecular dynamics simulations, and evaluation of Coulombic matrix elements similarly limits quantum mechanical self consistent field calculations. These difficulties are a result of the Coulomb operator’s slow decay, which…
Madushanka Manathunga, Hasan Metin Aktulga, Andreas W. Goetz, Kenneth M. Merz + 1 more
We have ported and optimized the GPU accelerated QUICK and AMBER based ab initio QM/MM implementation on AMD GPUs. This encompasses the entire Fock matrix build and force calculation in QUICK including one-electron integrals, two-electron repulsion integrals, exchange-correlation quadrature, and linear algebra…
Christopher Myers, Ken Miyazaki, Thomas Trepl, Christine Isborn + 1 more
GPU-accelerated on-the-fly nonadiabatic dynamics is enabled by interfacing the linearized semiclassical dynamics approach with the TeraChem electronic structure program. We describe the computational workflow of the "PySCES" code interface, a Python code for semiclassical dynamics with on-the-fly electronic structure…