11 papers · ranked by Valyu relevance
Alessandro Petrini, Marco Mesiti, Max Schubach, Marco Frasca + 7 more
Several prediction problems in Computational Biology and Genomic Medicine are characterized by both big data as well as a high imbalance between examples to be learned, whereby positive examples can represent a tiny minority with respect to negative examples. For instance, deleterious or pathogenic variants are…
Pierre Carrier, Bill Long, Richard Walsh, Jef Dawson + 4 more
High Performance Computing (HPC) Best Practice offers opportunities to implement lessons learned in areas such as computational chemistry and physics in genomics workflows, specifically Next-Generation Sequencing (NGS) workflows. In this study we will briefly describe how distributed-memory parallelism can be an…
Sikao Guo, Nenad Korolija, Kent Milfeld, Adip Jhaveri + 3 more
Particle-based reaction-diffusion models offer a high-resolution alternative to the continuum reaction-diffusion approach, capturing the discrete and volume-excluding nature of molecules undergoing stochastic dynamics. These methods are thus uniquely capable of simulating explicit self-assembly of particles into…
Pau Andrio, Adam Hospital, Cristian Ramon-Cortes, Javier Conejero + 4 more
The usage of workflows has led to progress in many fields of science, where the need to process large amounts of data is coupled with difficulty in accessing and efficiently using High Performance Computing platforms. On the one hand, scientists are focused on their problem and concerned with how to process their data.…
Wilfried Agbeto, Camille Coti, Vladimir Reinharz
Advances in graph algorithmics have allowed in-depth study of many natural objects from molecular biology or chemistry to social networks. Particularly in molecular biology and cheminformatics, understanding complex structures by identifying conserved sub-structures is a key milestone towards the artificial design of…
Stuart Byma, Akash Dhasade, Adrian Altenhoff, Christophe Dessimoz + 1 more
This paper presents a new, parallel implementation of clustering and demonstrates its utility in greatly speeding up the process of identifying homologous proteins. Clustering is a technique to reduce the number of comparison needed to find similar pairs in a set of n elements such as protein sequences. Precise…
Junichiro Makino, Toshikazu Ebisuzaki, Ryutaro Himeno, Yoshihide Hayashizaki
Rapidly increasing amount of short read data generated by NGSs (new-generation sequencers) calls for the development of fast and accurate read alignment programs. The programs based on hash table (BLAST) and Burrows-Wheeler transform (bwa-mem) are used, and the latter is known to give superior performance. We here…
Christoph Stelz, Lukas Hübner, Alexandros Stamatakis
Phylogenetic trees describe the evolutionary history among biological species based on their genomic data. Maximum Likelihood (ML) based phylogenetic inference tools search for the tree and evolutionary model that best explain the observed genomic data. Given the independence of likelihood score calculations between…
Kecong Tang, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang
Durbin’s positional Burrows-Wheeler transform (PBWT) enables algorithms with the optimal time complexity of O(MN) for reporting all vs all haplotype matches in a population panel with M haplotypes and N variant sites. However, even this efficiency may still be too slow when the number of haplotypes reaches millions. To…
J Kyle Medley, Shaik Asifullah, Joseph Hellerstein, Herbert M Sauro
Mechanistic kinetic models of biological pathways are an important tool for understanding biological systems. Constructing kinetic models requires fitting the parameters to experimental data. However, parameter fitting on these models is a non–convex, non–linear optimization problem. Many algorithms have been proposed…
Constantin Scholl, Kassian Kobert, Tomáš Flouri, Alexandros Stamatakis
Motivated by load balance issues in parallel calculations of the phylogenetic likelihood function, we recently introduced an approximation algorithm for efficiently distributing partitioned alignment data to a given number of CPUs. The goal is to balance the accumulated number of sites per CPU, and, at the same time…