14 papers · ranked by Valyu relevance
Dawei Shen, Yao-zhong Zhang, Seiya Imoto
Whole Slide Images (WSIs) are gigapixel, high-resolution digital scans of microscope slides, providing detailed tissue profiles for pathological analysis. Due to their gigapixel size and lack of detailed annotations, Multiple Instance Learning (MIL) becomes the primary technique for WSI analysis. However, current MIL…
Hui Zheng, Hai-Teng Wang, Wei-Bang Jiang, Zhong-Tao Chen + 5 more
While invasive brain-computer interfaces have shown promise for high-performance speech decoding under medical use, the potential of intracranial stereoElectroEn-cephaloGraphy (sEEG), which causes less damage to patients, remains underex-plored. With the rapid progress in representation learning, leveraging abundant…
Mahdi Pourmirzaei, Alex Morehead, Farzaneh Esmaili, Jarett Ren + 2 more
Converting protein tertiary structure into discrete tokens via vector-quantized variational autoencoders (VQ-VAEs) creates a language of 3D geometry and provides a natural interface between sequence and structure models. While pose invariance is commonly enforced, retaining chirality and directional cues without…
Mahdi Pourmirzaei, Alex Morehead, Farzaneh Esmaili, Jarett Ren + 2 more
Converting protein tertiary structure into discrete tokens via vector-quantized variational autoencoders (VQ-VAEs) creates a language of 3D geometry and provides a natural interface between sequence and structure models. While pose invariance is commonly enforced, retaining chirality and directional cues without…
Marika Kaden, Katrin Sophie Bohnsack, Mirko Weber, Mateusz Kudła + 3 more
We present an approach to investigate SARS-CoV-2 virus sequences based on alignment-free methods for RNA sequence comparison. In particular, we verify a given clustering result for the GISAID data set, which was obtained analyzing the molecular differences in coronavirus populations by phylogenetic trees. For this…
Julia Abel, Marika Kaden, Katrin Sophie Bohnsack, Mirko Weber + 2 more
In this contribution the discrimination between native and mirror models of proteins according to their chirality is tackled based on the structural protein information. This information is contained in the Ramachandran plots of the protein models. We provide an approach to classify those plots by means of an…
Yusri Dwi Heryanto, Yao-zhong Zhang, Seiya Imoto
Cell-type annotation in single-cell data involves identifying and labeling the cell types based on their gene expression profiles or molecular features. Recently, with advances in single-cell foundation models (FMs), unsupervised annotation and transfer learning with FMs have been explored for cell-type annotation…
Yufeng Liu, Linghui Chen, Haiyan Liu
The power of diffusion probabilistic models (DDPMs) in protein design was recently demonstrated by methods that performs three-dimensional protein backbone denoising. However, these DDPMs tend to generate protein backbones of idealized secondary structures and short loops, lacking diverse, non-idealized local…
Weiyi Xiao, Hegang Chen, Adrien Osakwe, Qihuang Zhang + 1 more
Spatial transcriptomic (ST) technologies enable the measurement of gene expression directly within tissue sections while preserving spatial context. Many ST platforms additionally generate paired histological images alongside spatially resolved transcriptomic profiles. However, most existing computational approaches…
Zhangyang Gao, Cheng Tan, Stan Z. Li
The equivariant nature of 3D coordinates has posed long term challenges in protein structure representation learning, alignment, and generation. Can we create a compact and invariant language that equivariantly represents protein structures? Towards this goal, we propose FoldToken2 to transfer equivariant structures…
Chao Pan, S. M. Hossein Tabatabaei Yazdi, S Kasra Tabatabaei, Alvaro G. Hernandez + 2 more
The main obstacles for the practical deployment of DNA-based data storage platforms are the prohibitively high cost of synthetic DNA and the large number of errors introduced during synthesis. In particular, synthetic DNA products contain both individual oligo (fragment) symbol errors as well as missing DNA oligo…
Qing Shao
We benchmark six numerical precision configurations for ESM-2 protein language models across throughput, memory footprint and predictive accuracy, on two workloads with sharply different characteristics: bulk embedding extraction and deep mutational scanning (DMS) variant-effect scoring. Accuracy is evaluated on the…
H. Robert Frost
We present an approach for modeling single cell RNA-sequencing (scRNA-seq) data using quaternions. Quaternions are four dimensional hypercomplex numbers that, along with real numbers, complex numbers and octonions, represent one of the four normed division algebras. Quaternions have been most widely employed to…
Chengting Yu, Yujie Wu, Aili Wang, Wolfgang Maass
Hyperdimensional computing (HDC) addresses massively parallel implementations of symbolic computations that are both more transparent than ANNs and LLMs and more suitable for in-memory computing on highly energy-efficient analog hardware. It captures an essential aspects of brain computations: objects, concepts, and…