11 papers · ranked by Valyu relevance
G.R. van der Ploeg, J.A. Westerhuis, A. Heintz-Buschart, A.K. Smilde
Recently, studies that investigate microbial temporal dynamics have become more frequent. In a longitudinal microbiome study design, microbial abundance data are collected across multiple time points from the same subjects. In this context, exploratory analysis of longitudinal microbiome data using Principal Component…
Abhinav Sharma, Davi Josué Marcon, Johannes Loubser, Karla Valéria Batista Lima + 2 more
The MTBseq pipeline, published in 2018, was designed to address bioinformatics challenges in tuberculosis research using whole-genome sequencing data. It was the first publicly available pipeline on Github to perform full analysis of whole-genome sequencing (WGS) data for Mycobacterium tuberculosis encompassing quality…
Kevin McDonnell, Nathan Wamsley, Jason Derks, Sarah Sipe + 3 more
The throughput of mass spectrometry (MS) proteomics can be increased substantially by multiplexing that enables parallelization of data acquisition. Such parallelization in the mass domain (plexDIA) and the time domain (timePlex) increases the density of mass spectra and the overlap between ions originating from…
Marissa E. Powers, Keith Mannthey, Priyanka Sebastian, Snehal Adsule + 6 more
Next Generation Sequencing (NGS) workloads largely consist of pipelines of tasks with heterogeneous compute, memory, and storage requirements. Identifying the optimal system configuration has historically required expertise in both system architecture and bioinformatics. This paper outlines infrastructure…
Andreas Härer, Diana J. Rennison
Parallel evolution of phenotypic traits is regarded as strong evidence for natural selection and has been studied extensively in a variety of taxa. However, we have limited knowledge of whether parallel evolution of host organisms is accompanied by parallel changes of their associated microbial communities (i.e.…
Kecong Tang, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang
Durbin’s positional Burrows-Wheeler transform (PBWT) enables algorithms with the optimal time complexity of O(MN) for reporting all vs all haplotype matches in a population panel with M haplotypes and N variant sites. However, even this efficiency may still be too slow when the number of haplotypes reaches millions. To…
Peiyu Zong, Wenpeng Deng, Jian Liu, Jue Ruan
The rapid advancements in sequencing length necessitate the adoption of increasingly efficient sequence alignment algorithms. The Needleman-Wunsch method introduces the foundational dynamic programming (DP) matrix calculation for global alignment, which evaluates the overall alignment of sequences. However, this method…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
KuaiKuai Duan, Rogers F. Silva, Md Abdur Rahaman, Zening Fu + 5 more
Multimodal data collected by international and national biobanking efforts have distinct scales and model orders and provide unique and complementary insights into disease mechanisms. We propose a novel, flexible and efficient data fusion approach—aNy-way independent component analysis (aNy-way ICA). aNy-way ICA fuses…
Sage Hahn, Max M. Owens, DeKang Yuan, Anthony C Juliano + 3 more
The use of pre-defined parcellations on surface-based representations of the brain as a method for data reduction is common across neuroimaging studies. In particular, prediction-based studies typically employ parcellation-driven summaries of brain measures as input to predictive algorithms, but the choice of…
Philipp S. L. Schäfer, Leoni Zimmermann, Paul L. Burmedi, Avia Walfisch + 7 more
Trade-offs between different functions or tasks are pervasive across scales in biological systems. For example, individual cells cannot perform all possible functions simultaneously; instead they allocate limited resources to specialize in subsets of tasks by activating specific gene expression programs. Pareto Task…