13 papers · ranked by Valyu relevance
Alejandro Hernandez Wences, Michael C. Schatz
Genome assembly projects typically run multiple algorithms in an attempt to find the single best assembly, although those assemblies often have complementary, if untapped, strengths and weaknesses. We present our metassembler algorithm that merges multiple assemblies of a genome into a single superior sequence. We…
Nikolas Bernaola, Mario Michiels, Pedro Larrañaga, Concha Bielza
We present the Fast Greedy Equivalence Search (FGES)-Merge, a new method for learning the structure of gene regulatory networks via merging locally learned Bayesian networks, based on the fast greedy equivalent search algorithm. The method is competitive with the state of the art in terms of the Matthews correlation…
Harun Mustafa, André Kahles, Mikhail Karasikov, Gunnar Rätsch
Much of the DNA and RNA sequencing data available is in the form of high-throughput sequencing (HTS) reads and is currently unindexed by established sequence search databases. Recent succinct data structures for indexing both reference sequences and HTS data, along with associated metadata, have been based on either…
Aaron S. Brewster, Daniel W. Paley, Asmit Bhowmick, David W. Mittan-Moreau + 5 more
The cctbx.xfel suite of processing programs and tools allows fast, visual analysis of serial diffraction images from synchrotrons and XFELs. Built on DIALS and cctbx, cctbx.xfel is designed for real-time and post-experiment processing with a fully featured graphical user interface. Users can quickly identify hitrates…
Noah Brown, Charles Danis, Vazira Ahmedjanova, Jennifer L. Guler
Structural variants (SVs) are abundant across all life, and have major impacts on the genome and transcriptome. However, it is difficult to appreciate the individual significance of SVs when they are heterogeneously distributed across a genomic neighborhood. Further, low-input sequencing technologies or sequencing of…
Thomas Krannich, W. Timothy J. White, Sebastian Niehus, Guillaume Holley + 2 more
With the increasing throughput of sequencing technologies, structural vari-ant (SV) detection has become possible across ten of thousands of genomes. Non-reference sequence (NRS) variants have drawn less attention compared to other types of SVs due to the computational complexity of detecting them. When using…
Vladimir Smirnov, Tandy Warnow
Phylogeny estimation is an important part of much biological research, but large-scale tree estimation is infeasible using standard methods due to computational issues. Recently, an approach to large-scale phylogeny has been proposed that divides a set of species into disjoint subsets, computes trees on the subsets…
Haleema Sadia, Sahal Sabilil Muttaqin, Parvez Alam
The accurate representation of color is important in applications involving species identification. Environmental variations introduce inconsistencies in color perception, affecting the reliability of automated image processing algorithms. In previous work, we developed a hybrid algorithm, AInsectID Version 1.1 Color…
Lisa Gandy, Jordan Gumm, Benjamin Fertig, Michael J. Kennish + 6 more
Today’s low cost digital data provides unprecedented opportunities for scientific discovery from synthesis studies. For example, the medical field is revolutionizing patient care by creating personalized treatment plans based upon mining electronic medical records, imaging, and genomics data. Standardized annotations…
Ksenia Khelik, Alexander Johan Nederbragt, Geir Kjetil Sandve, Torbjørn Rognes
In spite of the major breakthroughs in the second-generation sequencing technologies and the developments of a plethora of assemblers over the last ten years, the resulting genome assemblies may still be fragmented and contain errors. It is typical in genome projects with second-generation reads involved to run…
Adam C. English, Vipin K. Menon, Richard Gibbs, Ginger A. Metcalf + 1 more
For multi-sample structural variant analyses like merging, benchmarking, and annotation, the fundamental operation is to identify when two SVs are the same. Commonly applied approaches for comparing SVs were developed alongside technologies which produce ill-defined boundaries. As SV detection becomes more exact…
Vladimir Smirnov
Multiple sequence alignment tools struggle to keep pace with rapidly growing sequence data, as few methods can handle large datasets while maintaining alignment accuracy. We recently introduced MAGUS, a new state-of-the-art method for aligning large numbers of sequences. In this paper, we present a comprehensive set of…
Philip M. Hubbard, Stuart Berg, Ting Zhao, Donald J. Olbris + 6 more
Recent advances in automatic image segmentation and synapse prediction in electron microscopy (EM) datasets of the brain enable more efficient reconstruction of neural connectivity. In these datasets, a single neuron can span thousands of images containing complex tree-like arbors with thousands of synapses. While…