13 papers · ranked by Valyu relevance
Tsolak Ghukasyan, Vahagn Altunyan, Aram Bughdaryan, Tigran Aghajanyan + 3 more
This paper presents the Smart Distributed Data Factory (SDDF), an AI-driven distributed computing platform designed to address challenges in drug discovery by creating comprehensive datasets of molecular conformations and their properties. SDDF uses volunteer computing, leveraging the processing power of personal…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Haotian Li
Machine learning and deep learning are novel and trending approaches to solving real-world scientific problems. Graph machine learning is dedicated to performing learning methods, such as graph neural networks, on non-Euclidean data such as graphs. Molecules, with their natural graph structures, could be analyzed by…
Cláudia Brito, Pedro Ferreira, João Paulo
Breakthroughs in sequencing technologies led to an exponential growth of genomic data, providing unprecedented biological in-sights and new therapeutic applications. However, analyzing such large amounts of sensitive data raises key concerns regarding data privacy, specifically when the information is outsourced to…
Jared Streich, Anna Furches, David Kainer, Benjamin J. Garcia + 7 more
We present an exascale approach for producing global scale, high resolution, longitudinally based geoclimate classifications. Using a GPU implementation of the DUO Similarity Metric on the Summit supercomputer, we calculated the pairwise environmental similarity of 156,384,190 vectors of 414,640 encoded elements…
Christiam Camacho, Grzegorz M. Boratyn, Victor Joukov, Roberto Vera Alvarez + 1 more
Biomedical researchers use alignments produced by BLAST (Basic Local Alignment Search Tool) to categorize their query sequences. Producing such alignments is an essential bioinformatics task that is well suited for the cloud. The cloud can perform many calculations quickly as well as store and access large volumes of…
Tanveer Ahmad, Chengxin Ma, Zaid Al-Ars, H. Peter Hofstee
Current cluster scaled genomics data processing solutions rely on big data frameworks like Apache Spark, Hadoop and HDFS for data scheduling, processing and storage. These frameworks come with additional computation and memory overheads by default. It has been observed that scaling genomics dataset processing beyond 32…
Christoph Stelz, Lukas Hübner, Alexandros Stamatakis
Phylogenetic trees describe the evolutionary history among biological species based on their genomic data. Maximum Likelihood (ML) based phylogenetic inference tools search for the tree and evolutionary model that best explain the observed genomic data. Given the independence of likelihood score calculations between…
Maxim Lippeveld, Daniel Peralta, Andrew Filby, Yvan Saeys
Due to high resolution and throughput of modern image cytometry platforms, morphologically profiling generated datasets poses a significant computational challenge. Here, we present Scalable Cytometry Image Processing (SCIP), an image processing software aimed at running on distributed high performance computing…
Konrad Rokicki, David Schauder, Donald J. Olbris, Cristian Goina + 8 more
HortaCloud is a cloud-based, open-source platform designed to facilitate the collaborative reconstruction of long-range projection neurons from whole-brain light microscopy data. By providing virtual environments directly within the cloud, it eliminates the need for costly and time-consuming data downloads, allowing…
Swier Garst, Julian Dekker, Marcel Reinders
Federated learning is an upcoming machine learning paradigm which allows data from multiple sources to be used for training of classifiers without the data leaving the source it originally resides. This can be highly valuable for use cases such as medical research, where gathering data at a central location can be…
Patrick McKeever, Varun Mittal, Bryce Fukuda, Ka Yee Yeung + 1 more
The exponential growth of omics data requires novel strategies for storage, transfer, and processing of said data. We present a scheduler based on the Temporal.io workflow framework which enables two key optimizations of bioinformatics workflows. Firstly, we enable users to transparently map workflow steps to diverse…
Kevin Kang, Jinwen Wo, Jon Jiang, Zhong Wang
We propose Adaptive Container Service (ACS), a new paradigm for deploying bioinformatics workflows in cloud computing environments. By encapsulating the entire workflow within a single virtual container, combined with automatic workflow checkpointing and dynamic migration to appropriately scaled containers, ACS-based…