12 papers · ranked by Valyu relevance
Marcin Cieślik, Cameron Mura
PaPy, which stands for parallel pipelines in Python, is a highly flexible framework that enables the construction of robust, scalable workflows for either generating or processing voluminous datasets. A workflow is created from user-written Python functions (nodes) connected by 'pipes' (edges) into a directed acyclic…
Jiasi Shen, Martin Rinard, Nikos Vasilakis
We present KumQuat, a system for automatically generating data parallel implementations of Unix shell commands and pipelines. The generated parallel versions split input streams, execute multiple instantiations of the original pipeline commands to process the splits in parallel, then combine the resulting parallel…
Houming Wu, Ling Chen, Wenjie Yu
Large Models Training Authors: ['Houming Wu' 'Ling Chen' 'Wenjie Yu'] With the increasing scale of models, the need for efficient distributed training has become increasingly urgent. Recently, many synchronous pipeline parallelism approaches have been proposed to improve training throughput. However, these approaches…
Nikos Vasilakis, Κωνσταντίνος Καλλάς, Konstantinos Mamouras, Achilles Benetopoulos + 1 more
'Achilles Benetopoulos' 'Lazar Cvetković'] This paper presents PaSh, a system for parallelizing POSIX shell scripts. Given a script, PaSh converts it to a dataflow graph, performs a series of semantics-preserving program transformations that expose parallelism, and then converts the dataflow graph back into a…
Shivam Handa, Κωνσταντίνος Καλλάς, Nikos Vasilakis, Martin Rinard
We present a dataflow model for modelling parallel Unix shell pipelines. To accurately capture the semantics of complex Unix pipelines, the dataflow model is order-aware, i.e., the order in which a node in the dataflow graph consumes inputs from different edges plays a central role in the semantics of the computation…
Daniele Dall’Olio, Nico Curti, Eugenio Fonzi, Claudia Sala + 3 more
'Daniel Remondini' 'Gastone Castellani' 'Enrico Giampieri'] Background Current high-throughput technologies-i.e. whole genome sequencing, RNA-Seq, ChIP-Seq, etc.-generate huge amounts of data and their usage gets more widespread with each passing year. Complex analysis pipelines involving several…
Juan Cabral, B. Sánchez, M. Beroiz, M. Domínguez + 3 more
'S. Gurovich' 'Pablo M. Granitto'] Data processing pipelines represent an important slice of the astronomical software library that include chains of processes that transform raw data into valuable information via data reduction and analysis. In this work we present Corral, a Python framework for astronomical pipeline…
Marco Aldinucci, Cristina Calcagno, Mario Coppo, Ferruccio Damiani + 5 more
'Maurizio Drocco' 'Eva Sciacca' 'Salvatore Spinella' 'Massimo Torquati' 'Angelo Troina'] The paper arguments are on enabling methodologies for the design of a fully parallel, online, interactive tool aiming to support the bioinformatics scientists .In particular, the features of these methodologies, supported by the…
Marcin Cieślik, Cameron Mura
Background Bioinformatic analyses typically proceed as chains of data-processing tasks. A pipeline, or 'workflow', is a well-defined protocol, with a specific structure defined by the topology of data-flow interdependencies, and a particular functionality arising from the data transformations applied at each step. In…
Michael Grauer, Patrick Reynolds, Marion Hoogstoel, Francois Budin + 2 more
'Martin A. Styner' 'Ipek Oguz'] Image processing is an important quantitative technique for neuroscience researchers, but difficult for those who lack experience in the field. In this paper we present a web-based platform that allows an expert to create a brain image processing pipeline, enabling execution of that…
Satoshi Ito, Masaaki Yadome, Tatsuo Nishiki, Shigeru Ishiduki + 3 more
'Hikaru Inoue' 'Rui Yamaguchi' 'Satoru Miyano'] Background Supercomputers have become indispensable infrastructures in science and industries. In particular, most state-of-the-art scientific results utilize massively parallel supercomputers ranked in TOP500. However, their use is still limited in the bioinformatics…
Onur Yukselen, Osman Turkyilmaz, Ahmet Rasit Ozturk, Manuel Garber + 1 more
'Alper Kucukural'] Background The emergence of high throughput technologies that produce vast amounts of genomic data, such as next-generation sequencing (NGS) is transforming biological research. The dramatic increase in the volume of data, the variety and continuous change of data processing tools, algorithms and…