Optimized distributed systems achieve significant performance improvement on sorted merging of massive VCF files
Xiaobo Sun, Jingjing Gao, Peng Jin, Celeste Eng, Esteban G Burchard, Terri H Beaty, Ingo Ruczinski, Rasika A Mathias, Kathleen Barnes, Fusheng Wang, Zhaohui S Qin
Abstract
With the rapid development of high-throughput biotechnologies, genetic studies have entered the Big Data era. Studies like genome-wide association studies (GWASs), whole-genome sequencing (WGS), and whole-exome sequencing studies have produced massive amounts of data. The ability to efficiently manage and process such massive amounts of data becomes increasingly important for successful large-scale genetics studies []. Single-machine based methods are inefficient when processing such large amounts of data due to the prohibitive computation time, Input/Output bottleneck, as well as central proc

§ The Valyu brief
Reading the full paper and taking notes. This takes a few seconds…
§ Ask this paper
Ask a question about this paper
Valyu reads the full text and answers from what the paper actually says.
Searching the other archives…