23 papers · ranked by Valyu relevance
Cheng-Hsiang Chiu, Tsung‐Wei Huang, Zizheng Guo, Yibo Lin
Pipeline is a fundamental parallel programming pattern. Mainstream pipeline programming frameworks count on data abstractions to perform pipeline scheduling. This design is convenient for data-centric pipeline applications but inefficient for algorithms that only exploit task parallelism in pipeline. As a result, we…
Houming Wu, Ling Chen, Wenjie Yu
Large Models Training Authors: ['Houming Wu' 'Ling Chen' 'Wenjie Yu'] With the increasing scale of models, the need for efficient distributed training has become increasingly urgent. Recently, many synchronous pipeline parallelism approaches have been proposed to improve training throughput. However, these approaches…
Jesper Larsson Träff
These lecture notes are designed to accompany an imaginary, virtual, undergraduate, one or two semester course on fundamentals of Parallel Computing as well as to serve as background and reference for graduate courses on High-Performance Computing, parallel algorithms and shared-memory multiprocessor programming. They…
Marissa E. Powers, Keith Mannthey, Priyanka Sebastian, Snehal Adsule + 6 more
Next Generation Sequencing (NGS) workloads largely consist of pipelines of tasks with heterogeneous compute, memory, and storage requirements. Identifying the optimal system configuration has historically required expertise in both system architecture and bioinformatics. This paper outlines infrastructure…
Branden Butler, Sixing Yu, Arya Mazaheri, Ali Jannesari
Speculation Authors: ['Branden Butler' 'Sixing Yu' 'Arya Mazaheri' 'Ali Jannesari'] Abstract—Inference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from CPU speculative execution. These techniques…
Marek Sztuka, Krzysztof Kotlarz, Magda Mielczarek, Piotr Hajduk + 2 more
This study compared computational approaches to parallelisation of an SNP calling workflow. Data comprised DNA from five Holstein-Friesian cows sequenced with the Illumina platform. The pipeline consisted of quality control, alignment to the reference genome, post-alignment, and SNP calling. Three approaches to…
Ye Tian, Zhen Jia, Ziyue Luo, Yida Wang + 1 more
Diffusion models have emerged as dominant performers for image generation. To support training large diffusion models, this paper studies pipeline parallel training of diffusion models and proposes DiffusionPipe, a synchronous pipeline training system that advocates innovative pipeline bubble filling technique…
Jaroslav Budiš, Werner Krampl, Marcel Kucharík, Rastislav Hekel + 12 more
'Adrián Goga' 'Jozef Sitarčík' 'Michal Lichvár' 'Dávid Smol’ak' 'Miroslav Böhmer' 'Andrej Baláž' 'František Ďuriš' 'Juraj Gazdarica' 'Katarína Šoltys' 'Ján Turňa' 'Ján Radvánszky' 'Tomáš Szemes'] Title: Abstract With the rapid growth of massively parallel sequencing technologies, still more laboratories are utilising…
Jihu Guo, T.P. Ma, Wei Gao, Peng Sun + 4 more
Pipeline parallelism is widely used to train large language models (LLMs). However, increasing heterogeneity in model architectures exacerbates pipeline bubbles, thereby reducing training efficiency. Existing approaches overlook the cooptimization of model partition, model placement, and workload scheduling, resulting…
Marek Sztuka, Krzysztof Kotlarz, Magda Mielczarek, Piotr Hajduk + 2 more
'Jakub Liu' 'Joanna Szyda'] Title: Abstract This study compared computational approaches to parallelization of an SNP calling workflow. The data comprised DNA from five Holstein-Friesian cows sequenced with the Illumina platform. The pipeline consisted of quality control, alignment to the reference genome…
Edgardo M. Ortiz, Alina Höwener, Gentaro Shigita, Mustafa Raza + 5 more
A diverse range of high-throughput sequencing data, such as target capture, RNA-Seq, genome skimming, and high-depth whole genome sequencing, are used for phylogenomic analyses but the integration of such mixed data types into a single phylogenomic dataset requires a number of bioinformatic tools and significant…
Edgardo M. Ortiz, Alina Höwener, Gentaro Shigita, Mustafa Raza + 5 more
A diverse range of high-throughput sequencing data, such as target capture, RNA-Seq, genome skimming, and high-depth whole genome sequencing, are amenable to phylogenomic analyses but the integration of such mixed data types into a single phylogenomic dataset requires a number of bioinformatic tools and significant…
Jochen Sieg, Christian Wolfgang Feldmann, Jennifer Hemmerich, Conrad Stork + 3 more
The open-source package scikit-learn provides various machine learning algorithms and data processing tools, including the Pipeline class, which allows users to prepend custom data transformation steps to the machine learning model. We introduce the MolPipeline package, which extends this concept to chemoinformatics by…
Olesya Melnichenko, Venkat S. Malladi
In the field of genomics, bioinformatics pipelines play a crucial role in processing and analyzing vast biological datasets. These pipelines, consisting of interconnected tasks, can be optimized for efficiency and scalability by leveraging cloud platforms such as Microsoft Azure. The choice of compute resources…
Ashish Chapagain, Dima Abuoliem, In Ho Cho, Tongbiao Wang
Multifunctional nanosurfaces receive growing attention due to their versatile properties. Capillary force lithography (CFL) has emerged as a simple and economical method for fabricating these surfaces. In recent works, the authors proposed to leverage the evolution strategies (ES) to modify nanosurface characteristics…
Paul Cardosi, Bérenger Bramas, Bilal Alatas
Parallelization is needed everywhere, from laptops and mobile phones to supercomputers. Among parallel programming models, task-based programming has demonstrated a powerful potential and is widely used in high-performance scientific computing. Not only does it allow efficient parallelization across distributed…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…
Authors not listed
Protein conformational landscapes contain the functionally relevant information useful for understanding biological processes. Mapping out conformational landscapes provides valuable insights into protein behaviors and biological phenomena, and has relevance to therapeutic design. While experimental structural biology…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Peter Kraus, Edan Bainglass, Francisco F. Ramirez, Enea Svaluto-Ferro + 7 more
Compliance with good research data management practices means trust in the integrity of the data, and it is achievable by a full control of the data gathering process. In this work, we demonstrate tooling which bridges these two aspects, and illustrate its use in a case study of automated battery cycling. We…
Károly Bósa, Paul Heinzlreiter
Background Data preparation is a fundamental aspect of data engineering, a prerequisite for later tasks such as data visualization, reporting, and training machine learning models. Despite the recurring patterns in data transformation processes, the specific steps often vary depending on the project context, data…
Raúl Miñón, Josu Diaz-de-Arcaya, Ana I. Torre-Bastida, Juan López-de-Armentia + 4 more
'Juan López-de-Armentia' 'Gorka Zarate' 'Lander Bonilla' 'Asier Garcia-Perez' 'Jon Aguirre-Usandizaga'] Machine learning is already integrated in diverse domains enhancing their performance and decision support. For laboratories, this approach is normally sufficient. However, in real environments, these models can not…