20 papers · ranked by Valyu relevance
Benjamin Paaßen, Irena Koprinska, Kalina Yacef
Tree data occurs in many forms, such as computer programs, chemical molecules, or natural language. Unfortunately, the non-vectorial and discrete nature of trees makes it challenging to construct functions with tree-formed output, complicating tasks such as optimization or time series prediction. Autoencoders address…
David Harry Richman, Cheng Zhang, Frederick A. Matsen IV
As part of work to connect phylogenetics with machine learning, there has been considerable recent interest in vector encodings of phylogenetic trees. We present a simple new “ordered leaf attachment” (OLA) method for uniquely encoding a binary, rooted phylogenetic tree topology as an integer vector. OLA encoding and…
Amirmohammad Farzaneh, Mihai-Alin Badiu, Justin P. Coon
—Measuring the complexity of tree structures can be beneficial in areas that use tree data structures for storage, communication, and processing purposes. This complexity can then be used to compress tree data structures to their informationtheoretic limit. Additionally, the lack of models for random generation of…
Suzanne J Matthews, Tiffani L Williams
Background Biologists require new algorithms to efficiently compress and store their large collections of phylogenetic trees. Our previous work showed that TreeZip is a promising approach for compressing phylogenetic trees. In this paper, we extend our TreeZip algorithm by handling trees with weighted branches.…
Jie Zhang, En‐hui Yang, John C. Kieffer
We consider the problem of lossless compression of binary trees, with the aim of reducing the number of code bits needed to store or transmit such trees. A lossless grammar-based code is presented which encodes each binary tree into a binary codeword in two steps. In the first step, the tree is transformed into a…
Amirmohammad Farzaneh, Mihai-Alin Badiu, Justin P. Coon
—Routing tables in ad hoc and wireless routing protocols can be represented using rooted trees. The constant need for communication and storage of these trees in routing protocols demands an efficient rooted tree coding algorithm. This efficiency is defined in terms of the average code length, and the optimality of the…
Cedric Chauve, Caroline Colijn, Louxin Zhang
Good representations for phylogenetic trees and networks are important for enhancing storage efficiency and scalability for the inference and analysis of evolutionary trees for genes, genomes and species. We propose a new representation for rooted phylogenetic trees that encodes a tree on $n$ ordered taxa as a vector…
Christian Steinruecken
—This article describes lossless compression algorithms for multisets of sequences, taking advantage of the multiset's unordered structure. Multisets are a generalisation of sets where members are allowed to occur multiple times. A multiset can be encoded na¨ıvely by simply storing its elements in some sequential…
Axel Trefzer, Alexandros Stamatakis
Bayesian Markov-Chain Monte Carlo (MCMC) methods for phylogenetic tree inference, that is, inference of the evolutionary history of distinct species using their molecular sequence data, typically generate large sets of phylogenetic trees. The trees generated by the MCMC procedure are samples of the posterior…
Lukas Hübner, Alexandros Stamatakis
The field of population genetics attempts to advance our understanding of evolutionary processes. It has applications, for example, in medical research, wildlife conservation, and – in conjunction with recent advances in ancient DNA sequencing technology – studying human migration patterns over the past few thousand…
Carl Kingsford, Rob Patro
Motivation: Storing, transmitting and archiving data produced by next-generation sequencing is a significant computational burden. New compression techniques tailored to short-read sequence data are needed. Results: We present here an approach to compression that reduces the difficulty of managing large-scale…
W Timothy J White, Michael D Hendy
Background Publicly available DNA sequence databases such as GenBank are large, and are growing at an exponential rate. The sheer volume of data being dealt with presents serious storage and data communications problems. Currently, sequence data is usually kept in large "flat files," which are then compressed using…
Helen L. Smith, Patrick J. Biggs, Nigel P. French, Adam N. H. Smith + 2 more
'Jonathan C. Marshall' 'Muhammad Aleem'] Performance of random forest classification models is often assessed and interpreted using out-of-bag (OOB) samples. Observations which are OOB when a tree is trained may serve as a test set for that tree and predictions from the OOB observations used to calculate OOB error and…
Vladimir Kondratyev, Marian Dryzhakov, Timur Gimadiev, Dmitriy Slutskiy
In this work, we provide further development of the junction tree variational autoencoder (JT VAE) architecture in terms of implementation and application of the internal feature space of the model. Pretraining of JT VAE on a large dataset and further optimization with a regression model led to a latent space that can…
Richard Apodaca
Despite its widespread use, Simplified Molecular Input Line Entry System (SMILES) remains underspecified. The lack of a detailed specification encourages improvisation by software developers, complicates data standardization efforts, and undermines extension development. Balsa, a reformulation of SMILES, addresses…
Andrew Simmonett, Bernard Brooks, Thomas Darden
Evaluation of noncovalent electrostatic interactions is the dominant bottleneck in classical molecular dynamics simulations, and evaluation of Coulombic matrix elements similarly limits quantum mechanical self consistent field calculations. These difficulties are a result of the Coulomb operator’s slow decay, which…
Jia Lu, Ryan Tsoi, Nan Luo, Yuanchi Ha + 8 more
Dynamical systems often generate distinct outputs according to different initial conditions, and one can infer the corresponding input configuration given an output. This property captures the essence of information encoding and decoding. Here, we demonstrate the use of self-organized patterns, combined with machine…
Anas Al-okaily, Abdelghani Tbakhi
Data compression is a challenging and increasingly important problem. As the amount of data generated daily continues to increase, efficient transmission and storage has never been more critical. In this study, a novel encoding algorithm is proposed, motivated by the compression of DNA data and associated…
Authors not listed
We present a simple yet efficient random (brute-force) algorithm for constructing solvated molecular systems. By placing solvent molecules at random positions and orientations within a simulation box, we circumvent the complexities typically associated with more sophisticated packing algorithms. The main computational…
Ian Holmes
We describe a strategy for constructing codes for DNA-based information storage by serial composition of weighted finite-state transducers. The resulting state machines can integrate correction of substitution errors; synchronization by interleaving watermark and periodic marker signals; conversion from binary to…