18 papers · ranked by Valyu relevance
David Harry Richman, Cheng Zhang, Frederick A. Matsen IV
As part of work to connect phylogenetics with machine learning, there has been considerable recent interest in vector encodings of phylogenetic trees. We present a simple new “ordered leaf attachment” (OLA) method for uniquely encoding a binary, rooted phylogenetic tree topology as an integer vector. OLA encoding and…
Paolo Bresolin, Fabio Vandin, Vera Pancaldi
A crucial preliminary step to consider when adopting a GNN is the choice of an effective encoding for the nodes of the input graphs. Since we want our model to extract embeddings also for unseen phylogenetic trees, we need an encoding that works for any $T\inT$, with alterations in $M$. However, a first problem is that…
Pengyu Liu, Mariel Vázquez, Nataša Jonoska
Tree structures appear in many fields of the life sciences, including phylogenetics, developmental biology and nucleic acid structures. Trees can be used to represent RNA secondary structures, which directly relate to the function of non-coding RNAs. Recent developments in sequencing technology and artificial…
David Linus Ostby
We introduce Stingy Context, a hierarchical tree-based compression scheme achieving 18:1 reduction in LLM context tokens for auto-coding tasks. Using our TREEFRAG exploit decomposition, we reduce a real source code base of 239k tokens to 11k tokens while preserving task fidelity. Empirical results across 12 Frontier…
Xinru Zhang, Shizhe Ding, Chungong Yu, Jianquan Zhao + 2 more
Accurate phylogenetic inference is crucial for understanding evolutionary relationships among species. Deep learning technique has been introduced for phylogenetic inference; however, the existing deep learning-based approaches either suffer from limited accuracy as they split inference into several disjoint stages, or…
Marcin Zukowski
Huffman encoding has been an enduring technique for 70+ years, ubiquitous in compression algorithms since its invention. In this paper we propose a new approach to Huffman coding, based on a data structure from wavelet trees. The resulting pivot-coded Huffman (PivCo-Huffman) enables high-performance SIMD-friendly…
Xinyuanmeng Yao, Xiao Ma
This paper first presents a new approach to evaluating the descriptive complexity of finite-length binary sequences. Specifically, we investigate the sequence-wise recovery behavior induced by polar compression and successive cancellation decoding (SCD), and define the polar complexity of a sequence as the minimum…
H. Yamamoto, Ken-ichi Iwata
This paper proposes a new lossless data compression coding scheme named an asymmetric encoding-decoding scheme (AEDS), which can be considered as a generalization of tANS (tabled variant of asymmetric numeral systems). In the AEDS, a data sequence s = s1s 2 · · · s n is encoded in backward order st, t = n, · · · , 2…
Peer Rheinboldt, Frédéric Berdoz, Roger Wattenhofer
One-shot block drafters for speculative decoding generate the full draft in a single forward pass, achieving strong throughput by eliminating sequential token generation. However, they predict each draft token conditioned only on the prefix context, with no dependence on previously drafted tokens. This…
Authors not listed
RNA molecules fold into complex three-dimensional structures that determine their function. A wide range of mathematical frameworks, such as chord diagrams, fatgraphs, and context-free grammars, have been used to represent these structures; however, these models have largely been developed from mathematical motivations…
Authors not listed
Phase equilibrium calculations are crucial in chemical engineering design and optimization processes. The PC-SAFT equation of state (EoS) can precisely calculate phase equilibrium, but is relatively complex and computationally intensive. Surrogate models are mathematically simple models that map or regress the…
Authors not listed
Generating novel, drug-like molecules with realistic synthetic pathways is an essential goal in computer-aided drug discovery, yet generative models often lack synthesis awareness, resulting in compounds that are difficult or impossible to produce. To overcome this limitation, models must optimize not only molecular…
Fabio Cumbo, Kabir Dhillon, Jayadev Joshi, Davide Chicco + 2 more
Viral species classification is crucial for understanding viral evolution, epidemiology, and developing effective diagnostics and treatments. Traditional methods often rely on sequence similarity, which can be challenging for rapidly evolving viruses. Pangenomes, offering a comprehensive representation of species’…
Adrian Tkachenko, Sepehr Salem, Ayotomiwa Ezekiel Adeniyi, Zülal Bingöl + 6 more
High-throughput sequencing (HTS) enables population-scale genomics but generates massive datasets, creating bottlenecks in storage, transfer, and analysis. FASTQ, the standard format for over two decades, stores one byte per base and one byte per quality score, leading to inefficient I/O, high storage costs, and…
Matthieu Vilain, Stéphane Aris-Brosou
The ever-growing amount of available biological data leads modern analysis to be performed on large datasets. Unfortunately, bioinformatics tools for preprocessing and analyzing data are not always designed to treat such large amounts of data efficiently. Notably, this is the case when encoding DNA and RNA sequences…
Yuhang Wang, Weihua Chen, Linjing Song, Zhiping Xu + 6 more
With the rapid growth of data volume in sensor networks, lossy source coding systems achieve high-efficiency data compression with low distortion under limited transmission bandwidth. However, conventional compression algorithms rely on a two-stage framework with high computational complexity and frequently struggle to…
Yann Collet, Nick Terrell, W. Felix Handte, Danielle Rozenblit + 9 more
In the last few decades, research techniques have improved lossless compression ratios by significantly increasing processing time. However, these techniques have not gained popularity in industry because production systems require high throughput and low resource utilization. Instead, real world improvements in…
Authors not listed
Accurate prediction of chemical reaction yields remains essential for accelerating synthesis optimization, yet current machine learning models face critical limitations in capturing temporal dynamics, providing calibrated uncertainty estimates, and explicitly modeling reactant-to-product transformations. Here we…