12 papers · ranked by Valyu relevance
Marcin Cieślik, Cameron Mura
PaPy, which stands for parallel pipelines in Python, is a highly flexible framework that enables the construction of robust, scalable workflows for either generating or processing voluminous datasets. A workflow is created from user-written Python functions (nodes) connected by 'pipes' (edges) into a directed acyclic…
Jiasi Shen, Martin Rinard, Nikos Vasilakis
We present KumQuat, a system for automatically generating data parallel implementations of Unix shell commands and pipelines. The generated parallel versions split input streams, execute multiple instantiations of the original pipeline commands to process the splits in parallel, then combine the resulting parallel…
Houming Wu, Ling Chen, Wenjie Yu
Large Models Training Authors: ['Houming Wu' 'Ling Chen' 'Wenjie Yu'] With the increasing scale of models, the need for efficient distributed training has become increasingly urgent. Recently, many synchronous pipeline parallelism approaches have been proposed to improve training throughput. However, these approaches…
Nikos Vasilakis, Κωνσταντίνος Καλλάς, Konstantinos Mamouras, Achilles Benetopoulos + 1 more
'Achilles Benetopoulos' 'Lazar Cvetković'] This paper presents PaSh, a system for parallelizing POSIX shell scripts. Given a script, PaSh converts it to a dataflow graph, performs a series of semantics-preserving program transformations that expose parallelism, and then converts the dataflow graph back into a…
Shivam Handa, Κωνσταντίνος Καλλάς, Nikos Vasilakis, Martin Rinard
We present a dataflow model for modelling parallel Unix shell pipelines. To accurately capture the semantics of complex Unix pipelines, the dataflow model is order-aware, i.e., the order in which a node in the dataflow graph consumes inputs from different edges plays a central role in the semantics of the computation…
Abhinav Sharma, Davi Josué Marcon, Johannes Loubser, Karla Valéria Batista Lima + 2 more
The MTBseq pipeline, published in 2018, was designed to address bioinformatics challenges in tuberculosis research using whole-genome sequencing data. It was the first publicly available pipeline on Github to perform full analysis of whole-genome sequencing (WGS) data for Mycobacterium tuberculosis encompassing quality…
Juan Cabral, B. Sánchez, M. Beroiz, M. Domínguez + 3 more
'S. Gurovich' 'Pablo M. Granitto'] Data processing pipelines represent an important slice of the astronomical software library that include chains of processes that transform raw data into valuable information via data reduction and analysis. In this work we present Corral, a Python framework for astronomical pipeline…
Pau Andrio, Adam Hospital, Cristian Ramon-Cortes, Javier Conejero + 4 more
The usage of workflows has led to progress in many fields of science, where the need to process large amounts of data is coupled with difficulty in accessing and efficiently using High Performance Computing platforms. On the one hand, scientists are focused on their problem and concerned with how to process their data.…
Marek Sztuka, Krzysztof Kotlarz, Magda Mielczarek, Piotr Hajduk + 2 more
This study compared computational approaches to parallelisation of an SNP calling workflow. Data comprised DNA from five Holstein-Friesian cows sequenced with the Illumina platform. The pipeline consisted of quality control, alignment to the reference genome, post-alignment, and SNP calling. Three approaches to…
Dimitri Desvillechabrol, Rachel Legendre, Claire Rioualen, Christiane Bouchier + 3 more
We designed a PyQt graphical user interface – Sequanix – aiming at democratizing the use of Snakemake pipelines. Although the primary goal of Sequanix was to facilitate the execution of NGS Snakemake pipelines available in the Sequana project (http://sequana.readthedocs.io), it can also handle any Snakemake pipelines.…
Pierre Carrier, Bill Long, Richard Walsh, Jef Dawson + 4 more
High Performance Computing (HPC) Best Practice offers opportunities to implement lessons learned in areas such as computational chemistry and physics in genomics workflows, specifically Next-Generation Sequencing (NGS) workflows. In this study we will briefly describe how distributed-memory parallelism can be an…
Jason P. Kurs, Manuele Simi, Fabien Campagne
Computational workflows and pipelines are often created to automate series of processing steps. For instance, workflows enable one to standardize analysis for large projects or core facilities, but are also useful for individual biologists who need to perform repetitive data processing. Some workflow systems, designed…