13 papers · ranked by Valyu relevance
Abhinav Sharma, Davi Josué Marcon, Johannes Loubser, Karla Valéria Batista Lima + 2 more
The MTBseq pipeline, published in 2018, was designed to address bioinformatics challenges in tuberculosis research using whole-genome sequencing data. It was the first publicly available pipeline on Github to perform full analysis of whole-genome sequencing (WGS) data for Mycobacterium tuberculosis encompassing quality…
Marissa E. Powers, Keith Mannthey, Priyanka Sebastian, Snehal Adsule + 6 more
Next Generation Sequencing (NGS) workloads largely consist of pipelines of tasks with heterogeneous compute, memory, and storage requirements. Identifying the optimal system configuration has historically required expertise in both system architecture and bioinformatics. This paper outlines infrastructure…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
Marek Sztuka, Krzysztof Kotlarz, Magda Mielczarek, Piotr Hajduk + 2 more
This study compared computational approaches to parallelisation of an SNP calling workflow. Data comprised DNA from five Holstein-Friesian cows sequenced with the Illumina platform. The pipeline consisted of quality control, alignment to the reference genome, post-alignment, and SNP calling. Three approaches to…
Benjamin C. Hitz, Jin-Wook Lee, Otto Jolanki, Meenakshi S. Kagda + 56 more
The Encyclopedia of DNA elements (ENCODE) project is a collaborative effort to create a comprehensive catalog of functional elements in the human genome. The current database comprises more than 19000 functional genomics experiments across more than 1000 cell lines and tissues using a wide array of experimental…
Andrew E. Webb, Scott W. Wolf, Ian M. Traniello, Sarah D. Kocher
The exponential growth in biological data generation has created an urgent need for efficient, reproducible computational analysis workflows. Here, we present pipemake, a computational platform designed to streamline the development and implementation of efficient and reproducible Snakemake workflows. pipemake creates…
Edgardo M. Ortiz, Alina Höwener, Gentaro Shigita, Mustafa Raza + 5 more
A diverse range of high-throughput sequencing data, such as target capture, RNA-Seq, genome skimming, and high-depth whole genome sequencing, are used for phylogenomic analyses but the integration of such mixed data types into a single phylogenomic dataset requires a number of bioinformatic tools and significant…
Edgardo M. Ortiz, Alina Höwener, Gentaro Shigita, Mustafa Raza + 5 more
A diverse range of high-throughput sequencing data, such as target capture, RNA-Seq, genome skimming, and high-depth whole genome sequencing, are amenable to phylogenomic analyses but the integration of such mixed data types into a single phylogenomic dataset requires a number of bioinformatic tools and significant…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Olesya Melnichenko, Venkat S. Malladi
In the field of genomics, bioinformatics pipelines play a crucial role in processing and analyzing vast biological datasets. These pipelines, consisting of interconnected tasks, can be optimized for efficiency and scalability by leveraging cloud platforms such as Microsoft Azure. The choice of compute resources…
Elizabeth Koning, Raga Krishnakumar
Generating phylogenetic trees from genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention in the literature, and has been significantly streamlined over the years. Given the volume of publicly available genetic data, obtaining genomes for a wide…
Tanveer Ahmad, Chengxin Ma, Zaid Al-Ars, H. Peter Hofstee
Current cluster scaled genomics data processing solutions rely on big data frameworks like Apache Spark, Hadoop and HDFS for data scheduling, processing and storage. These frameworks come with additional computation and memory overheads by default. It has been observed that scaling genomics dataset processing beyond 32…
Peiyu Zong, Wenpeng Deng, Jian Liu, Jue Ruan
The rapid advancements in sequencing length necessitate the adoption of increasingly efficient sequence alignment algorithms. The Needleman-Wunsch method introduces the foundational dynamic programming (DP) matrix calculation for global alignment, which evaluates the overall alignment of sequences. However, this method…