11 papers · ranked by Valyu relevance
Ben Langmead, Christopher Wilks, Valentin Antonescu, Rone Charles
General-purpose processors can now contain many dozens of processor cores and support hundreds of simultaneous threads of execution. To make best use of these threads, genomics software must contend with new and subtle computer architecture issues. We discuss some of these and propose methods for improving thread…
Hiruna Samarakoon, James M. Ferguson, Sasha P. Jenner, Timothy G. Amos + 3 more
Nanopore sequencing is an emerging technology that is being rapidly adopted in research and clinical genomics. We recently developed SLOW5, a new file format for storage and analysis of raw data from nanopore sequencing experiments. SLOW5 is a community-centric, open source format that offers considerable performance…
Ling-Hong Hung, Wes Lloyd, Radhika Agumbe Sridhar, Saranya Devi Athmalingam Ravishankar + 3 more
For many next-generation sequencing pipelines, the most computationally intensive step is the alignment of reads to a reference sequence. As a result, alignment software such as the Burrows-Wheeler Aligner (BWA) is optimized for speed and and is often executed in parallel on the cloud. However, there are other less…
Min Shuai, Xin Chen
Weighted gene co-expression network analysis (WGCNA) is an R package that can search highly related gene modules. The most time-consuming step of the whole analysis is to calculate the Topological Overlap Matrix (TOM) from the Adjacency Matrix in a single thread. This study changes it to multithreading. This paper uses…
Peiyu Zong, Wenpeng Deng, Jian Liu, Jue Ruan
The rapid advancements in sequencing length necessitate the adoption of increasingly efficient sequence alignment algorithms. The Needleman-Wunsch method introduces the foundational dynamic programming (DP) matrix calculation for global alignment, which evaluates the overall alignment of sequences. However, this method…
Pierre Carrier, Bill Long, Richard Walsh, Jef Dawson + 4 more
High Performance Computing (HPC) Best Practice offers opportunities to implement lessons learned in areas such as computational chemistry and physics in genomics workflows, specifically Next-Generation Sequencing (NGS) workflows. In this study we will briefly describe how distributed-memory parallelism can be an…
Tomasz Kowalski, Szymon Grabowski
FASTQ remains among the widely used formats for high-throughput sequencing data. Despite advances in specialized FASTQ compressors, they are still imperfect in terms of practical performance tradeoffs. We present a multi-threaded version of Pseudogenome-based Read Compressor (PgRC), an in-memory algorithm for…
Shumpei Morita, Jay T. Groves
T cells can recognize a few molecules of cognate antigen amongst vastly outnumbering non-cognate ligands. The T cell receptor (TCR) differentiates antigens based on antigen-TCR binding dwell time through a kinetic proofreading process. Historically, this has been modeled as the ligated receptor undergoing a series of…
Costanza Pascal, Herzeel Charlotte, Verachtert Wilfried
elPrep is an established multi-threaded framework for preparing SAM and BAM files in sequencing pipelines. To achieve good performance, its software architecture makes only a single pass through a SAM/BAM file for multiple preparation steps, and keeps sequencing data as much as possible in main memory. Similar to other…
Ying Zhou, Sharon R. Browning, Brian L. Browning
Segments of identity by descent (IBD) are used in many genetic analyses. We present a method for detecting identical-by-descent haplotype segments that is optimized for large-scale genotype data. Our method, called hap-IBD, combines a compressed representation of genotype data, the positional Burrows-Wheeler transform…
Xinwei Zhao, Eberhard Korsching
DNA and RNA nucleotide sequences are ubiquitous in all biological cells, serving as both a comprehensive library of capabilities for the cells and as an impressive regulatory system to control cellular function. The multi-alignment framework (MAF) provided in this study offers a user-friendly platform for sequence…