11 papers · ranked by Valyu relevance
Alessio P. Buccino, Arjun Sridhar, David Feng, Karel Svoboda + 1 more
The scale of in vivo electrophysiology has expanded in recent years, with simultaneous recordings across thousands of electrodes now becoming routine. These advances have enabled a wide range of discoveries, but they also impose substantial computational demands. Spike sorting, the procedure that extracts spikes from…
Maja Lehr, Mattea Unger, Tobias Abele, Stefan J. Maurer + 2 more
The development of sorting strategies that directly report on functional activity remains a bottleneck in synthetic cell research. Current methodologies typically rely on sequential label-dependent probing, which limits throughput. Here, we introduce a label-free, buoyancy-driven selection strategy in which the mode of…
Rahul Varki, Christina Boucher
Relative Lempel–Ziv (RLZ) is an effective compression method for large, repetitive collections; however, the fundamental primitives required to elevate it from a passive archival format to a tractable representation for compressed construction have yet to be fully established. In this paper, we introduce an algorithmic…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
Páll Melsted, Elís Mar Guðnýjarson, Jóhannes Nordal
We present a GPU implementation of kallisto for RNA-seq transcript quantification. By redesigning the core algorithms: pseudoalignment, equivalence class intersection, and the EM algorithm; for massively parallel execution on GPUs, we achieve a 30–50× speedup over multithreaded CPU kallisto. On a benchmark of 100…
Nhan Ly-Trong, Samuel Martin, Nick Goldman, Nicola De Maio + 1 more
Phylogenetic analysis is essential to genomic epidemiology, for example in tracing the origin and evolution of SARS-CoV-2 variants during the COVID-19 pandemic. We previously introduced CMAPLE, a single-threaded implementation of the MAPLE algorithm designed for large-scale epidemiological genomic datasets. CMAPLE can…
Zhejian Yu
Fast simulation of next-generation sequencing (NGS) data is vital for software development and benchmarking. Here we describe art_modern, an accelerated ART simulator that can simulate various NGS data. We accelerated ART using updated sampling algorithms, single-instruction multiple-data (SIMD) instruction-set…
Rick Beeloo, Ragnar Groot Koerkamp
Searching short DNA patterns such as barcodes, primers, or CRISPR spacers within sequencing reads or genomes is a fundamental task in bioinformatics. These problems are instances of multiple approximate string matching (MASM) [1], which requires locating all occurrences with up to k errors of multiple patterns of…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Etienne Conchon-Kerjan, Timothe Rouzé, Lucas Robidou, Florian Ingels + 1 more
Approximate membership query structures are used throughout sequence bioinformatics, from read screening and metagenomic classification to assembly, indexing, and error correction. Among them, Bloom filters remain the default choice. They are not the most efficient structures in either time or memory, but they provide…