13 papers · ranked by Valyu relevance
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Haotian Li
Machine learning and deep learning are novel and trending approaches to solving real-world scientific problems. Graph machine learning is dedicated to performing learning methods, such as graph neural networks, on non-Euclidean data such as graphs. Molecules, with their natural graph structures, could be analyzed by…
Maxim Lippeveld, Daniel Peralta, Andrew Filby, Yvan Saeys
Due to high resolution and throughput of modern image cytometry platforms, morphologically profiling generated datasets poses a significant computational challenge. Here, we present Scalable Cytometry Image Processing (SCIP), an image processing software aimed at running on distributed high performance computing…
Pierre Carrier, Bill Long, Richard Walsh, Jef Dawson + 4 more
High Performance Computing (HPC) Best Practice offers opportunities to implement lessons learned in areas such as computational chemistry and physics in genomics workflows, specifically Next-Generation Sequencing (NGS) workflows. In this study we will briefly describe how distributed-memory parallelism can be an…
Jamie Alnasir, Hugh P. Shanahan
The paper reviews the use of the Hadoop platform in Structural Bioinformatics applications. Specifically, we review a number of implementations using Hadoop of high-throughput analyses, e.g. ligand-protein docking and structural alignment, and their scalability in comparison with other batch schedulers and MPI. We find…
Patrick McKeever, Varun Mittal, Bryce Fukuda, Ka Yee Yeung + 1 more
The exponential growth of omics data requires novel strategies for storage, transfer, and processing of said data. We present a scheduler based on the Temporal.io workflow framework which enables two key optimizations of bioinformatics workflows. Firstly, we enable users to transparently map workflow steps to diverse…
Gaurav Kaushik, Sinisa Ivkovic, Janko Simonovic, Nebojsa Tijanic + 2 more
As biomedical data becomes increasingly easy to generate in large quantities, the methods used to analyze it have proliferated rapidly. However, for the insights gained from these analyses to be meaningful, the analysis methods themselves must be transparent and reproducible. To address this issue, numerous groups have…
Ben Blamey, Salman Toor, Martin Dahlö, Håkan Wieslander + 6 more
This paper introduces the HASTE Toolkit, a cloud-native software toolkit capable of partitioning data streams in order to prioritize usage of limited resources. This in turn enables more efficient data-intensive experiments. We propose a model that introduces automated, autonomous decision making in data pipelines…
Patrick Diep, Jose L. Cadavid, Alexander F. Yakunin, Alison P. McGuigan + 1 more
Protein purification is a ubiquitous operation in biochemistry and life sciences and represents a key step to producing purified proteins for research (understanding how proteins work) and various applications. The need for scalable and parallel protein purification systems keeps growing due to the increase in…
Jason P. Kurs, Manuele Simi, Fabien Campagne
Computational workflows and pipelines are often created to automate series of processing steps. For instance, workflows enable one to standardize analysis for large projects or core facilities, but are also useful for individual biologists who need to perform repetitive data processing. Some workflow systems, designed…
Moises Hernandez-Fernandez, Istvan Reguly, Saad Jbabdi, Mike Giles + 2 more
The great potential of computational diffusion MRI (dMRI) relies on indirect inference of tissue microstructure and brain connections, since modelling and tractography frameworks map diffusion measurements to neuroanatomical features. This mapping however can be computationally highly expensive, particularly given the…
Azza E Ahmed, Joshua M Allen, Tajesvi Bhat, Prakruthi Burra + 16 more
The changing landscape of genomics research and clinical practice has created a need for computational pipelines capable of efficiently orchestrating complex analysis stages while handling large volumes of data across heterogeneous computational environments. Workflow Management Systems (WfMSs) are the software…
Ben Langmead, Christopher Wilks, Valentin Antonescu, Rone Charles
General-purpose processors can now contain many dozens of processor cores and support hundreds of simultaneous threads of execution. To make best use of these threads, genomics software must contend with new and subtle computer architecture issues. We discuss some of these and propose methods for improving thread…