15 papers · ranked by Valyu relevance
Maria Clicia S. Castro, Vanessa S. Silva, Maiana O. C. Costa, Helena S. I. L. Silva + 8 more
Several hundred terabytes of single-cell RNA-seq (scRNA-seq) data are available in public repositories. These data refer to various research projects, from microbial population cells to multiple tissues, involving patients with a myriad of diseases and comorbidities. An increase to several Petabytes of these scRNA-seq…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Junichiro Makino, Toshikazu Ebisuzaki, Ryutaro Himeno, Yoshihide Hayashizaki
Rapidly increasing amount of short read data generated by NGSs (new-generation sequencers) calls for the development of fast and accurate read alignment programs. The programs based on hash table (BLAST) and Burrows-Wheeler transform (bwa-mem) are used, and the latter is known to give superior performance. We here…
Sebastian Schönherr, Johanna Schachtl-Riess, Silvia Di Maio, Michele Filosi + 5 more
Genome-wide association studies (GWAS) in large biobanks are transforming genetic research and enable the detection of novel genotype-phenotype relationships. In the last two decades, over 60,000 genetic associations across thousands of human diseases and traits have been discovered using a GWAS approach. Due to denser…
Chuanfang Ning, Gabriel Bunke, Simon Lietar, Lukas van den Heuvel + 1 more
We developed a versatile lab-on-chip (LOC) workcell that enables the design and automatic execution of experiments on LOC devices, improving how we establish, optimize, and productionalize LOC processes. Key features include direct docking and cooling of native laboratory tubes, programmable reagent mixing and…
Peiyu Zong, Wenpeng Deng, Jian Liu, Jue Ruan
The rapid advancements in sequencing length necessitate the adoption of increasingly efficient sequence alignment algorithms. The Needleman-Wunsch method introduces the foundational dynamic programming (DP) matrix calculation for global alignment, which evaluates the overall alignment of sequences. However, this method…
Sikao Guo, Nenad Korolija, Kent Milfeld, Adip Jhaveri + 3 more
Particle-based reaction-diffusion models offer a high-resolution alternative to the continuum reaction-diffusion approach, capturing the discrete and volume-excluding nature of molecules undergoing stochastic dynamics. These methods are thus uniquely capable of simulating explicit self-assembly of particles into…
Kecong Tang, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang
Durbin’s positional Burrows-Wheeler transform (PBWT) enables algorithms with the optimal time complexity of O(MN) for reporting all vs all haplotype matches in a population panel with M haplotypes and N variant sites. However, even this efficiency may still be too slow when the number of haplotypes reaches millions. To…
Wilfried Agbeto, Camille Coti, Vladimir Reinharz
Advances in graph algorithmics have allowed in-depth study of many natural objects from molecular biology or chemistry to social networks. Particularly in molecular biology and cheminformatics, understanding complex structures by identifying conserved sub-structures is a key milestone towards the artificial design of…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated quadratic…
Patrick McKeever, Varun Mittal, Bryce Fukuda, Ka Yee Yeung + 1 more
The exponential growth of omics data requires novel strategies for storage, transfer, and processing of said data. We present a scheduler based on the Temporal.io workflow framework which enables two key optimizations of bioinformatics workflows. Firstly, we enable users to transparently map workflow steps to diverse…
Stefano Marangoni, Federica Furia, Debora Charrance, Agata Fant + 9 more
Next Generation Sequencing (NGS) has revolutionized genome biology, enabling the rapid sequencing of an entire human genome and facilitating the integration of Whole Genome Sequencing (WGS) into both research and clinical applications. The high-throughput nature of NGS and the complex data processing required has…
Robin Henry, Jean-Baptiste Lugagne
Achieving real-time control of genetic systems is critical for improving the reliability, efficiency, and reproducibility of biological research and engineering. Yet the intrinsic stochasticity of these systems makes this goal difficult. Prior efforts have faced three recurring challenges: (a) predictive models of gene…