Search · four archives
Search · four archives
11 papers · ranked by Valyu relevance
Francesco Versaci, Luca Pireddu, Gianluigi Zanetti
Modern sequencing machines produce order of a terabyte of data per day, which need subsequently to go through a complex processing pipeline. The standard workflow begins with a few independent, shared-memory tools, which communicate by means of intermediate files. Given the constant increase of the amount of data…
Ben Blamey, Salman Toor, Martin Dahlö, Håkan Wieslander + 6 more
This paper introduces the HASTE Toolkit, a cloud-native software toolkit capable of partitioning data streams in order to prioritize usage of limited resources. This in turn enables more efficient data-intensive experiments. We propose a model that introduces automated, autonomous decision making in data pipelines…
Tanveer Ahmad, Chengxin Ma, Zaid Al-Ars, H. Peter Hofstee
Current cluster scaled genomics data processing solutions rely on big data frameworks like Apache Spark, Hadoop and HDFS for data scheduling, processing and storage. These frameworks come with additional computation and memory overheads by default. It has been observed that scaling genomics dataset processing beyond 32…
Peter G. Hawkins, Eli M. Swanson, Megan Feichtel
The size of individual single cell samples continues to grow with advancing technologies, as do the number of samples included in individual experiments and across organizations. This presents challenges for processing this data at scale, both in terms of computational throughput and the required size of the machines…
Francesco Versaci, Luca Pireddu, Gianluigi Zanetti
The adoption of Big Data technologies can potentially boost the scalability of data-driven biology and health workflows by orders of magnitude. Consider, for instance, that technologies in the Hadoop ecosystem have been successfully used in data-driven industry to scale their processes to levels much larger than any…
J.J. Salvo, M. Lakshman, A.M. Holubecki, Z.M. Saygin + 2 more
Reading bridges sensation and cognition. To derive meaning from written words, visual input is first processed in unimodal (i.e., sensory-specific) visual streams and then engages a distributed language network (LANG) that includes classic perisylvian language areas and supports transmodal (i.e., sensory-nonspecific)…
Binke Yuan, Hui Xie, Zhihao Wang, Yangwen Xu + 11 more
Modern linguistic theories and network science propose that the language and speech processing is organized into hierarchical, segregated large-scale subnetworks, with a core of dorsal (phonological) stream and ventral (semantic) stream. The two streams are asymmetrically recruited in receptive and expressive language…
Jon Ander Novella, Payam Emami Khoonsari, Stephanie Herman, Daniel Whitenack + 4 more
Computational biologists face many challenges related to data size, and they need to manage complicated analyses often including multiple stages and multiple tools, all of which must be deployed to modern infrastructures. To address these challenges and maintain reproducibility of results, researchers need (i) a…
Benjamin Coleman, Benito Geordie, Li Chou, R. A. Leo Elworth + 2 more
The rise of whole-genome shotgun sequencing (WGS) has enabled numerous breakthroughs in large-scale comparative genomics research. However, the size of genomic datasets has grown exponentially over the last few years, leading to new challenges for traditional streaming algorithms. Modern petabyte-sized genomic datasets…
Sergei Yakneen, Sebastian M. Waszak, Michael Gertz, Jan O. Korbel
We present Butler, a computational framework developed in the context of the international Pan-cancer Analysis of Whole Genomes (PCAWG)1 project to overcome the challenges of orchestrating analyses of thousands of human genomes on the cloud. Butler operates equally well on public and academic clouds. This highly…
Samuel Lampa, Martin Dahlö, Jonathan Alvarsson, Ola Spjuth
The complex nature of biological data has driven the development of specialized software tools. Scientific workflow management systems simplify the assembly of such tools into pipelines and assist with job automation and aids reproducibility of analyses. Many contemporary workflow tools are specialized and not designed…