13 papers · ranked by Valyu relevance
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Haotian Li
Machine learning and deep learning are novel and trending approaches to solving real-world scientific problems. Graph machine learning is dedicated to performing learning methods, such as graph neural networks, on non-Euclidean data such as graphs. Molecules, with their natural graph structures, could be analyzed by…
Maxim Lippeveld, Daniel Peralta, Andrew Filby, Yvan Saeys
Due to high resolution and throughput of modern image cytometry platforms, morphologically profiling generated datasets poses a significant computational challenge. Here, we present Scalable Cytometry Image Processing (SCIP), an image processing software aimed at running on distributed high performance computing…
Peter G. Hawkins, Eli M. Swanson, Megan Feichtel
The size of individual single cell samples continues to grow with advancing technologies, as do the number of samples included in individual experiments and across organizations. This presents challenges for processing this data at scale, both in terms of computational throughput and the required size of the machines…
Tanveer Ahmad, Chengxin Ma, Zaid Al-Ars, H. Peter Hofstee
Current cluster scaled genomics data processing solutions rely on big data frameworks like Apache Spark, Hadoop and HDFS for data scheduling, processing and storage. These frameworks come with additional computation and memory overheads by default. It has been observed that scaling genomics dataset processing beyond 32…
Wilfried Agbeto, Camille Coti, Vladimir Reinharz
Advances in graph algorithmics have allowed in-depth study of many natural objects from molecular biology or chemistry to social networks. Particularly in molecular biology and cheminformatics, understanding complex structures by identifying conserved sub-structures is a key milestone towards the artificial design of…
Patrick McKeever, Varun Mittal, Bryce Fukuda, Ka Yee Yeung + 1 more
The exponential growth of omics data requires novel strategies for storage, transfer, and processing of said data. We present a scheduler based on the Temporal.io workflow framework which enables two key optimizations of bioinformatics workflows. Firstly, we enable users to transparently map workflow steps to diverse…
Marissa E. Powers, Keith Mannthey, Priyanka Sebastian, Snehal Adsule + 6 more
Next Generation Sequencing (NGS) workloads largely consist of pipelines of tasks with heterogeneous compute, memory, and storage requirements. Identifying the optimal system configuration has historically required expertise in both system architecture and bioinformatics. This paper outlines infrastructure…
Patrick Diep, Jose L. Cadavid, Alexander F. Yakunin, Alison P. McGuigan + 1 more
Protein purification is a ubiquitous operation in biochemistry and life sciences and represents a key step to producing purified proteins for research (understanding how proteins work) and various applications. The need for scalable and parallel protein purification systems keeps growing due to the increase in…
Swier Garst, Julian Dekker, Marcel Reinders
Federated learning is an upcoming machine learning paradigm which allows data from multiple sources to be used for training of classifiers without the data leaving the source it originally resides. This can be highly valuable for use cases such as medical research, where gathering data at a central location can be…
David S. Cerutti, Rafal Wiewiora, Simon Boothroyd, Woody Sherman
The Structure and TOpology Replica Molecular Mechanics (STORMM) code is a next-generation molecular simulation engine and associated libraries optimized for performance on fast, multicore central processor units (CPUs) and graphics processing units (GPUs) with independent memory and tens of thousands of threads. STORMM…
Bertil Schmidt, Felix Kallenborn, Alejandro Chacon, Christian Hundt
The maximal sensitivity for local pairwise alignment makes the Smith-Waterman algorithm a popular choice for protein sequence database search. However, its quadratic time complexity makes it compute-intensive. Unfortunately, current state-of-the-art software tools are not able to leverage the massively parallel…
Kevin Kang, Jinwen Wo, Jon Jiang, Zhong Wang
We propose Adaptive Container Service (ACS), a new paradigm for deploying bioinformatics workflows in cloud computing environments. By encapsulating the entire workflow within a single virtual container, combined with automatic workflow checkpointing and dynamic migration to appropriately scaled containers, ACS-based…