15 papers · ranked by Valyu relevance
David Harry Richman, Cheng Zhang, Frederick A. Matsen IV
As part of work to connect phylogenetics with machine learning, there has been considerable recent interest in vector encodings of phylogenetic trees. We present a simple new “ordered leaf attachment” (OLA) method for uniquely encoding a binary, rooted phylogenetic tree topology as an integer vector. OLA encoding and…
Hanania, Dganit, Yaakobi, Eitan
—Labeling of DNA molecules is a fundamental technique for DNA visualization and analysis. This process was mathematically modeled in [1], where the received sequence indicates the positions of the used labels. In this work, we develop error correcting codes for labeled DNA sequences, establishing bounds and…
Florian L. Wagner, Gernot Neun, Thomas Tampone, Zhen Lei + 3 more
In this work we describe the development of a chemistry-based encoding approach utilizing nucleophilicity to perform Bayesian optimization campaigns. A fully automated slug continuous flow platform leveraging a liquid handler to investigate categorical variables is used for the self-optimization of organic reactions.…
Wenhao Liu, Zhengyi Jiang, Zhongyi Huang, Hanxu Hou
Fluorescent labeling is a cornerstone of DNA visualization and a key enabler of random access in DNA-based data storage. However, the stochastic nature of biochemical processes, including synthesis, hybridization, and optical readout, induces \emph{burst} synchronization errors within the resulting labeling sequences.…
Fabio Cumbo, Kabir Dhillon, Jayadev Joshi, Davide Chicco + 2 more
Viral species classification is crucial for understanding viral evolution, epidemiology, and developing effective diagnostics and treatments. Traditional methods often rely on sequence similarity, which can be challenging for rapidly evolving viruses. Pangenomes, offering a comprehensive representation of species’…
Li, Wangkai, Sun, Rui + 4 more
Pseudo-label learning is widely used in semantic segmentation, particularly in label-scarce scenarios such as unsupervised domain adaptation (UDA) and semisupervised learning (SSL). Despite its success, this paradigm can generate erroneous pseudo-labels, which are further amplified during training due to utilization of…
Marcel Nöhre, Gerd Stumme
We propose a flexible, two-phase algorithm for labeling line diagrams of ordered sets, in which the nodes of direct neighbors in the order relation are connected by a straight, upward-pointing line. In contrast to the labeling of diagrams of arbitrary graphs, we benefit from the fact that all edges in line diagrams of…
Igor Sadalski
Single-cell foundational models have emerged as a powerful tool for learning generalizable cellular representations from large-scale data. Most models in this domain use transformer backbones, which require careful engineering of gene and expression encoding strategies, yet there is no consensus on which encoding…
Yuhang Wang, Weihua Chen, Linjing Song, Zhiping Xu + 6 more
With the rapid growth of data volume in sensor networks, lossy source coding systems achieve high-efficiency data compression with low distortion under limited transmission bandwidth. However, conventional compression algorithms rely on a two-stage framework with high computational complexity and frequently struggle to…
Andrew Garrett Kurbis, Alex Mihailidis, Brokoslaw Laschowski
Decoding algorithms can be used to predict motor behaviour from patterns of neural activity. However, most studies rely on subject-optimized models, limiting generalization and scalability to novel subjects and tasks. Building on recent advances in deep learning and large-scale data, here we developed an EMG foundation…
Authors not listed
Conventional molecular graphs often are unable to reliably encode stereochemistry, especially for symmetric molecules, non-tetrahedral centers, and transition states. To overcome this, we present StereoMolGraph, an open source Python library implementing a stereochemistry-aware graph representation for molecules and…
Shuta Kikuchi, Shu Tanaka
The RNA inverse folding problem aims to identify nucleotide sequences that preferentially adopt a given target secondary structure. While various heuristic and machine learning-based approaches have been proposed, many require a large number of sequence evaluations, which limits their applicability when experimental…
Authors not listed
Olfaction arises from the interaction of odorants with olfactory receptors, a process shaped by molecular geometry, electron distribution, and conformational preference. We present ConfDENSE, a Set2Set enhanced PointNet model that learns directly from Hirshfeld promolecule electron-density point clouds, preserving full…
H. Yamamoto, Ken-ichi Iwata
This paper proposes a new lossless data compression coding scheme named an asymmetric encoding-decoding scheme (AEDS), which can be considered as a generalization of tANS (tabled variant of asymmetric numeral systems). In the AEDS, a data sequence s = s1s 2 · · · s n is encoded in backward order st, t = n, · · · , 2…
Rania Derouich, Nour El Houda Mathlouthi
We present the first systematic, hardware-executed benchmark of twelve distinct quantum data-encoding strategies for drug-response prediction on a real superconducting quantum processing unit (QPU). All experiments were conducted on the IQM Garnet 20-qubit QPU via the IQM Resonance cloud platform, using the Qrisp…