24 papers · ranked by Valyu relevance
Marcin Cieślik, Cameron Mura
PaPy, which stands for parallel pipelines in Python, is a highly flexible framework that enables the construction of robust, scalable workflows for either generating or processing voluminous datasets. A workflow is created from user-written Python functions (nodes) connected by 'pipes' (edges) into a directed acyclic…
Jiasi Shen, Martin Rinard, Nikos Vasilakis
We present KumQuat, a system for automatically generating data parallel implementations of Unix shell commands and pipelines. The generated parallel versions split input streams, execute multiple instantiations of the original pipeline commands to process the splits in parallel, then combine the resulting parallel…
Houming Wu, Ling Chen, Wenjie Yu
Large Models Training Authors: ['Houming Wu' 'Ling Chen' 'Wenjie Yu'] With the increasing scale of models, the need for efficient distributed training has become increasingly urgent. Recently, many synchronous pipeline parallelism approaches have been proposed to improve training throughput. However, these approaches…
Nikos Vasilakis, Κωνσταντίνος Καλλάς, Konstantinos Mamouras, Achilles Benetopoulos + 1 more
'Achilles Benetopoulos' 'Lazar Cvetković'] This paper presents PaSh, a system for parallelizing POSIX shell scripts. Given a script, PaSh converts it to a dataflow graph, performs a series of semantics-preserving program transformations that expose parallelism, and then converts the dataflow graph back into a…
Shivam Handa, Κωνσταντίνος Καλλάς, Nikos Vasilakis, Martin Rinard
We present a dataflow model for modelling parallel Unix shell pipelines. To accurately capture the semantics of complex Unix pipelines, the dataflow model is order-aware, i.e., the order in which a node in the dataflow graph consumes inputs from different edges plays a central role in the semantics of the computation…
Abhinav Sharma, Davi Josué Marcon, Johannes Loubser, Karla Valéria Batista Lima + 2 more
The MTBseq pipeline, published in 2018, was designed to address bioinformatics challenges in tuberculosis research using whole-genome sequencing data. It was the first publicly available pipeline on Github to perform full analysis of whole-genome sequencing (WGS) data for Mycobacterium tuberculosis encompassing quality…
Daniele Dall’Olio, Nico Curti, Eugenio Fonzi, Claudia Sala + 3 more
'Daniel Remondini' 'Gastone Castellani' 'Enrico Giampieri'] Background Current high-throughput technologies-i.e. whole genome sequencing, RNA-Seq, ChIP-Seq, etc.-generate huge amounts of data and their usage gets more widespread with each passing year. Complex analysis pipelines involving several…
Juan Cabral, B. Sánchez, M. Beroiz, M. Domínguez + 3 more
'S. Gurovich' 'Pablo M. Granitto'] Data processing pipelines represent an important slice of the astronomical software library that include chains of processes that transform raw data into valuable information via data reduction and analysis. In this work we present Corral, a Python framework for astronomical pipeline…
Marco Aldinucci, Cristina Calcagno, Mario Coppo, Ferruccio Damiani + 5 more
'Maurizio Drocco' 'Eva Sciacca' 'Salvatore Spinella' 'Massimo Torquati' 'Angelo Troina'] The paper arguments are on enabling methodologies for the design of a fully parallel, online, interactive tool aiming to support the bioinformatics scientists .In particular, the features of these methodologies, supported by the…
Marcin Cieślik, Cameron Mura
Background Bioinformatic analyses typically proceed as chains of data-processing tasks. A pipeline, or 'workflow', is a well-defined protocol, with a specific structure defined by the topology of data-flow interdependencies, and a particular functionality arising from the data transformations applied at each step. In…
Pau Andrio, Adam Hospital, Cristian Ramon-Cortes, Javier Conejero + 4 more
The usage of workflows has led to progress in many fields of science, where the need to process large amounts of data is coupled with difficulty in accessing and efficiently using High Performance Computing platforms. On the one hand, scientists are focused on their problem and concerned with how to process their data.…
Marek Sztuka, Krzysztof Kotlarz, Magda Mielczarek, Piotr Hajduk + 2 more
This study compared computational approaches to parallelisation of an SNP calling workflow. Data comprised DNA from five Holstein-Friesian cows sequenced with the Illumina platform. The pipeline consisted of quality control, alignment to the reference genome, post-alignment, and SNP calling. Three approaches to…
Michael Grauer, Patrick Reynolds, Marion Hoogstoel, Francois Budin + 2 more
'Martin A. Styner' 'Ipek Oguz'] Image processing is an important quantitative technique for neuroscience researchers, but difficult for those who lack experience in the field. In this paper we present a web-based platform that allows an expert to create a brain image processing pipeline, enabling execution of that…
Dimitri Desvillechabrol, Rachel Legendre, Claire Rioualen, Christiane Bouchier + 3 more
We designed a PyQt graphical user interface – Sequanix – aiming at democratizing the use of Snakemake pipelines. Although the primary goal of Sequanix was to facilitate the execution of NGS Snakemake pipelines available in the Sequana project (http://sequana.readthedocs.io), it can also handle any Snakemake pipelines.…
Jochen Sieg, Christian Wolfgang Feldmann, Jennifer Hemmerich, Conrad Stork + 3 more
The open-source package scikit-learn provides various machine learning algorithms and data processing tools, including the Pipeline class, which allows users to prepend custom data transformation steps to the machine learning model. We introduce the MolPipeline package, which extends this concept to chemoinformatics by…
Satoshi Ito, Masaaki Yadome, Tatsuo Nishiki, Shigeru Ishiduki + 3 more
'Hikaru Inoue' 'Rui Yamaguchi' 'Satoru Miyano'] Background Supercomputers have become indispensable infrastructures in science and industries. In particular, most state-of-the-art scientific results utilize massively parallel supercomputers ranked in TOP500. However, their use is still limited in the bioinformatics…
Pierre Carrier, Bill Long, Richard Walsh, Jef Dawson + 4 more
High Performance Computing (HPC) Best Practice offers opportunities to implement lessons learned in areas such as computational chemistry and physics in genomics workflows, specifically Next-Generation Sequencing (NGS) workflows. In this study we will briefly describe how distributed-memory parallelism can be an…
Jason P. Kurs, Manuele Simi, Fabien Campagne
Computational workflows and pipelines are often created to automate series of processing steps. For instance, workflows enable one to standardize analysis for large projects or core facilities, but are also useful for individual biologists who need to perform repetitive data processing. Some workflow systems, designed…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Onur Yukselen, Osman Turkyilmaz, Ahmet Rasit Ozturk, Manuel Garber + 1 more
'Alper Kucukural'] Background The emergence of high throughput technologies that produce vast amounts of genomic data, such as next-generation sequencing (NGS) is transforming biological research. The dramatic increase in the volume of data, the variety and continuous change of data processing tools, algorithms and…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…
Ido Ben-Shalom, Charles Lin, Brian Radak, Woody Sherman + 1 more
Molecular dynamics (MD) simulations of proteins are commonly used to sample from the Boltzmann distribution of conformational states, with wide-ranging applications spanning chemistry, biophysics, and drug discovery. However, MD can be inefficient at equilibrating water occupancy for buried cavities in proteins that…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Peter Kraus, Edan Bainglass, Francisco F. Ramirez, Enea Svaluto-Ferro + 7 more
Compliance with good research data management practices means trust in the integrity of the data, and it is achievable by a full control of the data gathering process. In this work, we demonstrate tooling which bridges these two aspects, and illustrate its use in a case study of automated battery cycling. We…