14 papers · ranked by Valyu relevance
Sumesh Kumar, Joseph Zambreno, Ashfaq Khokhar, Shoaib Akram + 1 more
Improving the speed and efficiency of database search algorithms that deduce peptides from mass spectrometry (MS) data has been an active area of research for more than three decades. The significance of the need for faster database search methods has rapidly increased due to the growing interest in studying non-model…
Nhan Ly-Trong, Samuel Martin, Nick Goldman, Nicola De Maio + 1 more
Phylogenetic analysis is essential to genomic epidemiology, for example in tracing the origin and evolution of SARS-CoV-2 variants during the COVID-19 pandemic. We previously introduced CMAPLE, a single-threaded implementation of the MAPLE algorithm designed for large-scale epidemiological genomic datasets. CMAPLE can…
Max Ward, Mary Richardson, Haining Lin, Michael Stamm + 7 more
mRNA medicines hold great promise, but designing sequences with high translation efficiency, robust in-solution stability, and manufacturability remains a major challenge due to the vast combinatorial space of synonymous coding sequences. Computational approaches such as mRNA folding algorithms have emerged as powerful…
Rick Beeloo, Ragnar Groot Koerkamp
Searching short DNA patterns such as barcodes, primers, or CRISPR spacers within sequencing reads or genomes is a fundamental task in bioinformatics. These problems are instances of multiple approximate string matching (MASM) [1], which requires locating all occurrences with up to k errors of multiple patterns of…
Rafael Terra, Diego Carvalho, Denis Jacob Machado, Carla Osthoff + 1 more
Advances in High-Performance Computing (HPC) have enabled increasingly complex genomic analyses, including those in phylogenomics. These analyses contribute to understanding the evolution of viruses and pathogens, improving our knowledge of disease transmission, and supporting targeted public health strategies.…
Ye Yuan, Zhe Li, Kaiqiang Hu, Pengwei Pan + 1 more
Head-to-tail (H2T) cyclized peptides are an increasingly important modality in drug discovery, combining high target affinity and selectivity with metabolic stability. Because their underlying chemistry is still that of a linear amino-acid chain, their linear sequence representation is the native input format of…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Páll Melsted, Elís Mar Guðnýjarson, Jóhannes Nordal
We present a GPU implementation of kallisto for RNA-seq transcript quantification. By redesigning the core algorithms: pseudoalignment, equivalence class intersection, and the EM algorithm; for massively parallel execution on GPUs, we achieve a 30–50× speedup over multithreaded CPU kallisto. On a benchmark of 100…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Noam Teyssier, Alexander Dobin
Single-cell genomics is rapidly scaling toward billion-cell atlases, but computational analysis has become a critical bottleneck. Processing multiplexed datasets with existing tools requires substantial computational resources and runtime that become prohibitive at scale. Here we present cyto, an ultra highthroughput…
Christos Papalitsas, Ioannis Mouratidis, Michail Patsakis, Evangelos Stogiannos + 2 more
The exponential growth of publicly available genomic data has created unprecedented opportunities for sequence-based discovery. Locating specific k-mers is fundamental to diverse applications, including metagenomic classification, pathogen and cancer detection, and variant calling yet efficient identification of…
Etienne Conchon-Kerjan, Timothe Rouzé, Lucas Robidou, Florian Ingels + 1 more
Approximate membership query structures are used throughout sequence bioinformatics, from read screening and metagenomic classification to assembly, indexing, and error correction. Among them, Bloom filters remain the default choice. They are not the most efficient structures in either time or memory, but they provide…
Haziq Moinudeen, Alp Duygu
Functional protein sequence space is often described either by the sparsity of functional sequences or by the connectivity of neutral networks. These quantities characterize different properties: density measures the fraction of all sequences that are functional, whereas connectivity describes relationships among…