13 papers · ranked by Valyu relevance
Joseph Moon, Peer-Timo Bremer, Pratik Mukherjee, Amy J. Markowitz + 5 more
Large scale diffusion MRI tractography remains a significant challenge. Users must orchestrate a complex sequence of instructions that require many software packages with complex dependencies and high computational cost. We developed MaPPeRTrac, a diffusion MRI tractography pipeline that simplifies and vastly…
Abhinav Sharma, Davi Josué Marcon, Johannes Loubser, Karla Valéria Batista Lima + 2 more
The MTBseq pipeline, published in 2018, was designed to address bioinformatics challenges in tuberculosis research using whole-genome sequencing data. It was the first publicly available pipeline on Github to perform full analysis of whole-genome sequencing (WGS) data for Mycobacterium tuberculosis encompassing quality…
Marissa E. Powers, Keith Mannthey, Priyanka Sebastian, Snehal Adsule + 6 more
Next Generation Sequencing (NGS) workloads largely consist of pipelines of tasks with heterogeneous compute, memory, and storage requirements. Identifying the optimal system configuration has historically required expertise in both system architecture and bioinformatics. This paper outlines infrastructure…
Pau Andrio, Adam Hospital, Cristian Ramon-Cortes, Javier Conejero + 4 more
The usage of workflows has led to progress in many fields of science, where the need to process large amounts of data is coupled with difficulty in accessing and efficiently using High Performance Computing platforms. On the one hand, scientists are focused on their problem and concerned with how to process their data.…
Edgardo M. Ortiz, Alina Höwener, Gentaro Shigita, Mustafa Raza + 5 more
A diverse range of high-throughput sequencing data, such as target capture, RNA-Seq, genome skimming, and high-depth whole genome sequencing, are used for phylogenomic analyses but the integration of such mixed data types into a single phylogenomic dataset requires a number of bioinformatic tools and significant…
Marek Sztuka, Krzysztof Kotlarz, Magda Mielczarek, Piotr Hajduk + 2 more
This study compared computational approaches to parallelisation of an SNP calling workflow. Data comprised DNA from five Holstein-Friesian cows sequenced with the Illumina platform. The pipeline consisted of quality control, alignment to the reference genome, post-alignment, and SNP calling. Three approaches to…
Dimitri Desvillechabrol, Rachel Legendre, Claire Rioualen, Christiane Bouchier + 3 more
We designed a PyQt graphical user interface – Sequanix – aiming at democratizing the use of Snakemake pipelines. Although the primary goal of Sequanix was to facilitate the execution of NGS Snakemake pipelines available in the Sequana project (http://sequana.readthedocs.io), it can also handle any Snakemake pipelines.…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Pierre Carrier, Bill Long, Richard Walsh, Jef Dawson + 4 more
High Performance Computing (HPC) Best Practice offers opportunities to implement lessons learned in areas such as computational chemistry and physics in genomics workflows, specifically Next-Generation Sequencing (NGS) workflows. In this study we will briefly describe how distributed-memory parallelism can be an…
Jason P. Kurs, Manuele Simi, Fabien Campagne
Computational workflows and pipelines are often created to automate series of processing steps. For instance, workflows enable one to standardize analysis for large projects or core facilities, but are also useful for individual biologists who need to perform repetitive data processing. Some workflow systems, designed…
Ben Langmead, Christopher Wilks, Valentin Antonescu, Rone Charles
General-purpose processors can now contain many dozens of processor cores and support hundreds of simultaneous threads of execution. To make best use of these threads, genomics software must contend with new and subtle computer architecture issues. We discuss some of these and propose methods for improving thread…
Moises Hernandez-Fernandez, Istvan Reguly, Saad Jbabdi, Mike Giles + 2 more
The great potential of computational diffusion MRI (dMRI) relies on indirect inference of tissue microstructure and brain connections, since modelling and tractography frameworks map diffusion measurements to neuroanatomical features. This mapping however can be computationally highly expensive, particularly given the…
Onur Yukselen, Osman Turkyilmaz, Ahmet Rasit Ozturk, Manuel Garber + 1 more
The emergence of high throughput technologies that produce vast amounts of genomic data, such as next-generation sequencing (NGS) are transforming biological research. The dramatic increase in the volume of data makes analysis the main bottleneck for scientific discovery. The processing of high throughput datasets…