12 papers · ranked by Valyu relevance
Nhan Ly-Trong, Samuel Martin, Nick Goldman, Nicola De Maio + 1 more
Phylogenetic analysis is essential to genomic epidemiology, for example in tracing the origin and evolution of SARS-CoV-2 variants during the COVID-19 pandemic. We previously introduced CMAPLE, a single-threaded implementation of the MAPLE algorithm designed for large-scale epidemiological genomic datasets. CMAPLE can…
Sumesh Kumar, Joseph Zambreno, Ashfaq Khokhar, Shoaib Akram + 1 more
Improving the speed and efficiency of database search algorithms that deduce peptides from mass spectrometry (MS) data has been an active area of research for more than three decades. The significance of the need for faster database search methods has rapidly increased due to the growing interest in studying non-model…
Max Ward, Mary Richardson, Haining Lin, Michael Stamm + 7 more
mRNA medicines hold great promise, but designing sequences with high translation efficiency, robust in-solution stability, and manufacturability remains a major challenge due to the vast combinatorial space of synonymous coding sequences. Computational approaches such as mRNA folding algorithms have emerged as powerful…
Rick Beeloo, Ragnar Groot Koerkamp
Searching short DNA patterns such as barcodes, primers, or CRISPR spacers within sequencing reads or genomes is a fundamental task in bioinformatics. These problems are instances of multiple approximate string matching (MASM) [1], which requires locating all occurrences with up to k errors of multiple patterns of…
Ye Yuan, Zhe Li, Kaiqiang Hu, Pengwei Pan + 1 more
Head-to-tail (H2T) cyclized peptides are an increasingly important modality in drug discovery, combining high target affinity and selectivity with metabolic stability. Because their underlying chemistry is still that of a linear amino-acid chain, their linear sequence representation is the native input format of…
Páll Melsted, Elís Mar Guðnýjarson, Jóhannes Nordal
We present a GPU implementation of kallisto for RNA-seq transcript quantification. By redesigning the core algorithms: pseudoalignment, equivalence class intersection, and the EM algorithm; for massively parallel execution on GPUs, we achieve a 30–50× speedup over multithreaded CPU kallisto. On a benchmark of 100…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Noam Teyssier, Alexander Dobin
Single-cell genomics is rapidly scaling toward billion-cell atlases, but computational analysis has become a critical bottleneck. Processing multiplexed datasets with existing tools requires substantial computational resources and runtime that become prohibitive at scale. Here we present cyto, an ultra highthroughput…
Mahmudur Rahman Hera, David Koslicki, Conrado Martínez
With the surge in sequencing data generated from an ever-expanding range of biological studies, designing scalable computational techniques has become essential. One effective strategy to enable large-scale computation is to split long DNA or protein sequences into k-mers, and summarize large k-mer sets into compact…
Florian Ingels, Léa Vandamme, Mathilde Girard, Clément Agret + 2 more
Modern sequencing continues to drive explosive growth of nucleotide sequence archives, pushing MinHash-derived sketching methods to their practical scalability limits. State-of-the-art tools such as Mash, Dashing2, and Bindash2 provide compact sketches and accurate similarity estimates for large collections, yet…