12 papers · ranked by Valyu relevance
Ling-Hong Hung, Niharika Nasam, Chris Biju, Wes Lloyd + 1 more
Singe cell RNA sequencing (scRNA-seq) has become a routine method for measuring cell activities. Processing large scRNA-seq datasets requires high-performance computing resources. The emergence of cloud computing allows us to leverage its on-demand capabilities without major investment in infrastructure. Serverless…
Ling-Hong Hung, Niharika Nasam, Wes Lloyd, Ka Yee Yeung
Singe cell RNA sequencing (scRNA-seq) has become a routine method for measuring cell activities. Processing large scRNA-seq datasets requires high-performance computing resources. The emergence of cloud computing allows us to leverage its on-demand capabilities without major investment in infrastructure. Serverless…
Ling-Hong Hung, Dimitar Kumanov, Xingzhi Niu, Wes Lloyd + 1 more
We have used serverless AWS Lambda functions to align 640 million reads in less than 3 minutes, a speed-up of 500x over the single-threaded implementation. Using a hybrid cloud architecture and software modified to optimize disk transfers, an entire RNA sequencing workflow transforming multiplexed reads to transcript…
Jody Clements, Cristian Goina, Philip M. Hubbard, Takashi Kawase + 4 more
Neuroscience research in Drosophila is benefiting from large-scale connectomics efforts using electron microscopy (EM) to reveal all the neurons in a brain and their connections. In order to exploit this knowledge base, researchers target individual neurons and study their function. Therefore, vast libraries of fly…
Andrea Kölzsch, Sarah C. Davidson, Dominik Gauggel, Clemens Hahn + 10 more
Bio-logging and animal tracking datasets continuously grow in volume and complexity, documenting animal behaviour and ecology in unprecedented extent and detail, but greatly increasing the challenge of extracting knowledge from the data obtained. A large variety of analysis methods are being developed, many of which in…
Jacob Bradford, Divya Joy, Mattias Winsen, Nicholas Meurant + 4 more
Gene editing has been revolutionised by the CRISPR-Cas9 technology. The versatility and ease-of-use of the technology far exceeds its predecessors, however, the selection of a high-quality guide RNA (gRNA) is critical to directing it to a target site. Selecting gRNA calls upon high-performance algorithms that evaluate…
Daniel P. Ramirez-Echemendia, Luís Borges-Araújo, Chelsea M. Brown, Riccardo Alessandri + 3 more
The Martini coarse-grained force field is widely used for biomolecular simulations by a large and rapidly expanding community worldwide. Over time, the development of Martini parameters, tools, and documentation has become increasingly dispersed across numerous research groups, leading to fragmentation and making it…
Soohyun Lee, Jeremy Johnson, Carl Vitzthum, Koray Kırlı + 2 more
We introduce Tibanna, an open-source software tool for automated execution of bioinformatics pipelines on Amazon Web Services (AWS). Tibanna accepts reproducible and portable pipeline standards including Common Workflow Language (CWL), Workflow Description Language (WDL) and Docker. It adopts a strategy of isolation…
Chih Chuan Shih, Jieqi Chen, Ai Shan Lee, Nicolas Bertin + 34 more
Genomic researchers are increasingly utilizing commercial cloud platforms (CCPs) to manage their data and analytics needs. Commercial clouds allow researchers to grow their storage and analytics capacity on demand, keeping pace with expanding project data footprints and enabling researchers to avoid large capital…
Jacob M. Luber, Braden T. Tierney, Evan M. Cofer, Chirag J. Patel + 1 more
Across biology we are seeing rapid developments in scale of data production without a corresponding increase in data analysis capabilities. Here, we present Aether (http://aether.kosticlab.org), an intuitive, easy-to-use, cost-effective, and scalable framework that uses linear programming (LP) to optimally bid on and…
Payam Emami Khoonsari, Pablo Moreno, Sven Bergmann, Joachim Burman + 31 more
Developing a robust and performant data analysis workflow that integrates all necessary components whilst still being able to scale over multiple compute nodes is a challenging task. We here present a generic method based on the microservice architecture, where software tools are encapsulated as Docker containers that…
PJ Tatlow, Stephen R. Piccolo
Public compendia of raw sequencing data are now measured in petabytes. Accordingly, it is becoming infeasible for individual researchers to transfer these data to local computers. Recently, the National Cancer Institute funded an initiative to explore opportunities and challenges of working with molecular data in…