13 papers · ranked by Valyu relevance
Sumesh Kumar, Joseph Zambreno, Ashfaq Khokhar, Shoaib Akram + 1 more
Improving the speed and efficiency of database search algorithms that deduce peptides from mass spectrometry (MS) data has been an active area of research for more than three decades. The significance of the need for faster database search methods has rapidly increased due to the growing interest in studying non-model…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Lin Chen, Yuhan Chen, Ziqi Cheng, Jing Guo + 20 more
MHPC512 is a massively parallel, special-purpose supercomputer designed primarily for atomic-level molecular dynamics (MD) simulations of biomolecular systems. It comprises 512 processor units interconnected by a high-speed three-dimensional torus network and employs a custom chip architecture that uses 35-bit…
Rafael Terra, Diego Carvalho, Denis Jacob Machado, Carla Osthoff + 1 more
Advances in High-Performance Computing (HPC) have enabled increasingly complex genomic analyses, including those in phylogenomics. These analyses contribute to understanding the evolution of viruses and pathogens, improving our knowledge of disease transmission, and supporting targeted public health strategies.…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
Marco Savioli, Paolo Calligari, Ugo Locatelli, Gianfranco Bocchinfuso
We introduce GROMODEX, a novel tool designed to optimise GROMACS molecular dynamics (MD) simulations using a structured Design of Experiments (DoE) approach. GROMACS, though efficient, requires extensive tuning of parameters to perform optimally on different hardware and molecular systems. Manual tuning is tedious and…
Páll Melsted, Elís Mar Guðnýjarson, Jóhannes Nordal
We present a GPU implementation of kallisto for RNA-seq transcript quantification. By redesigning the core algorithms: pseudoalignment, equivalence class intersection, and the EM algorithm; for massively parallel execution on GPUs, we achieve a 30–50× speedup over multithreaded CPU kallisto. On a benchmark of 100…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Nhan Ly-Trong, Samuel Martin, Nick Goldman, Nicola De Maio + 1 more
Phylogenetic analysis is essential to genomic epidemiology, for example in tracing the origin and evolution of SARS-CoV-2 variants during the COVID-19 pandemic. We previously introduced CMAPLE, a single-threaded implementation of the MAPLE algorithm designed for large-scale epidemiological genomic datasets. CMAPLE can…
Chun Gong, Qi Yang, Ruiwen Wan, Shengkang Li + 2 more
Joint variant calling is a crucial step in population-scale sequencing analysis. While population-scale sequencing is a powerful tool for genetic studies, achieving fast and accurate joint variant calling on large cohorts remains computationally challenging. To meet this challenge, we developed Distributed Population…
Emre Green, Adil Mardinoglu
Whole-genome sequencing (WGS) has transformed clinical diagnostics, yet variant annotation remains a computational bottleneck. The Variant Effect Predictor (VEP) integrates pathogenicity predictors and population databases essential for ACMG/AMP variant classification, but these annotation plugins are fundamentally…
Bonson Wong, Gagandeep Singh, Haris Javaid, Kristof Denolf + 4 more
Nanopore sequencing technologies are used widely in genomics research and their adoption continues to accelerate. ‘Basecalling’ is an essential step in the nanopore sequencing workflow, during which raw electrical signals are translated into nucleotide sequences. The current state-of-the-art basecaller, Oxford Nanopore…
Stephan Grein, David R. Penas, Daniel Weindl, Polina Lakrisenko + 2 more
Dynamic models are central to the computational life sciences but typically contain unknown parameters that must be inferred from experimental data. High-throughput measurements have made this task increasingly challenging, yielding high-dimensional search spaces and non-convex objectives with many local optima. This…