26 papers · ranked by Valyu relevance
Luke Squires, Jose Humberto Giraldez Chavez, Alfred Nilsson, Lukas Käll + 1 more
Better Learning: A Peptide Embedding Tutorial for Proteomic Mass Spectrometry Authors: Luke Squires, Jose Humberto Giraldez Chavez, Alfred Nilsson, Lukas Käll, Samuel H Payne Mass spectrometry proteomics creates complex data representing the peptide/protein contents of biological samples. Various types of machine…
Simchoni, Giora, Rosset, Saharon
We present MMbeddings, a probabilistic embedding approach that reinterprets categorical embeddings through the lens of nonlinear mixed models, effectively bridging classical statistical theory with modern deep learning. By treating embeddings as latent random effects within a variational autoencoder framework, our…
Kabane, Siyaxolisa
We investigate the generalization properties of dense text embeddings when the embedding backbone is a large language model (LLM) versus when it is a non-LLM encoder, and we study the extent to which spherical linear interpolation (SLERP) model-merging mitigates over-specialization introduced by task-specific…
Xin Yuan, Ke Chen, Ajmain Yasar Ahmed, Mingfu Shao
Edit distance is a fundamental metric for quantifying similarity between biological sequences, but its high computational cost limits large-scale applications. Previously, we proposed learned locality-sensitive bucketing (LSB) functions that achieved superior performance and efficiency compared to classical seeding…
Yuanchuan Guo, Jun S. Liu, Huimin Cheng, Ying Ma
As spatially resolved transcriptomics (SRT) datasets increasingly span multiple adjacent or replicated slices, effective joint analysis across slices is needed to reconstruct tissue structures and identify consistent spatial gene expression patterns. This requires resolving spatial correspondences between slices while…
Marta Bistroń, Jacek M. Żurada, Zbigniew Piotrowski, Euntai Kim + 2 more
Highlights What are the main findings?1. Deep learning-based watermarking methods (CNN, GAN, Transformers, and diffusion models) significantly outperform traditional spatial- and frequency-domain techniques in terms of robustness, transparency, and adaptability to modern attack types. 2. Emerging architectures such as…
Hsiang-Cheh Huang, Feng-Cheng Chang, Hong-Yi Li, Rabi N. Mahapatra + 1 more
With the proliferation of image-capturing and display-enabled IoT devices, ensuring the authenticity and integrity of visual data has become increasingly critical, especially in light of emerging cybersecurity threats and powerful generative AI tools. One of the major challenges in such sensor-based systems is the…
Dongnam Byun, Jungwon Park, Jungmin Ko, Changin Choi + 1 more
Background on CLIP text embeddings in text-toimage generative models. Let P be the input prompt for image generation, which is passed through CLIP text encoders to produce the CLIP text embedding: [ cSOT, c t 1 , . . . , c t L , cEOT, cPAD, . . . , cPAD ] ∈ R 77×d , where SOT, EOT, and PAD denote special tokens…
Navid NaderiAlizadeh, Rohit Singh
Protein language models (PLMs) encode amino acid sequences into residue-level embeddings that must be pooled into fixed-size representations for downstream protein-level prediction tasks. Although these embeddings implicitly reflect evolutionary constraints, existing pooling strategies operate on single sequences and…
Authors not listed
High-level quantum mechanical (QM) simulations provide accurate electronic information of chemical systems but scale unfavourably with system size, making calculations of applied systems challenging. Hierarchical quantum mechanics in quantum mechanics embedding (QM/QM) addresses this issue by localising the highly…
Haozhe Shan, Ashok Litwin-Kumar
In many circuit models of neural computation, synaptic connections between neurons are organized according to their tuning to the variables being processed. The connectivities of canonical neural network models of head direction, spatial navigation, and orientation selectivity obey this principle and contain symmetries…
Zeyu Xu, Xingzhong Hou, Pengkai Guo, Siling Lin + 5 more
Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Experts (MoE) models to improve performance limits deployment in resource-constrained environments due to the trade-off between high memory usage from full loading and increased latency…
Huabo Shen, Xiaodong Sun, Youmin Hu, Changgeng Li + 3 more
Zero-shot learning (ZSL) aims to classify unseen classes by leveraging semantic information from seen classes, addressing the challenge of limited labeled data. In recent years, ZSL methods have focused on extracting attribute-level features from images and aligning them with semantic features within an embedding…
Authors not listed
Channelrhodopsin-2 (ChR2) is a light-gated ion channel widely used in optogenetics, a technique that enables precise control of neuronal activity by genetically engineering light-sensitive proteins into cell membranes. This protein exists in dimeric form, with each monomer containing a retinal Schiff base (RSB) moiety…
Kris Sankaran, Shuzhen Zhang, Chenab, Marina Meilă
Nonlinear dimensionality reduction methods like Uniform Manifold Approximation and Projection (UMAP) and T-distributed stochastic neighbor embedding (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek}…
Harshita Sahni, Xin Chen, Trilce Estrada
Protein language models (PLMs) generate rich, layer-wise embeddings that capture diverse biological information but are expensive in terms of storage and computation at scale. In this work, we propose a compact surrogate representation for PLM embeddings across transformer layers using low-dimensional PCA projections…
Jiarong Lu, Bin Liao, Yi Liu, Lei Zhong
Against the complex characteristics of the Ethereum transaction network and the limitations of existing graph embedding methods based on random walks, which fail to effectively capture transaction temporal dynamics and the flow of funds, we propose a fraud detection algorithm for Ethereum, ETX2Vec (Ethereum…
Authors not listed
Recent years have seen a growing interest in machine learning approaches for chemical tasks. The best existing methods focus on building base models that combine molecular graphs (“2D structures”) with atomic coordinates in 3D to predict molecular properties, typically through pre-training followed by fine-tuning on…
Joseph Fioresi, Fabian Caba Heilbron, Pankaj Nathani, Mubarak Shah + 1 more
Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification. These embeddings compress high-level semantics into a single vector, which comes at the cost of primarily expressing a dominant semantics like main object while…
Li, Jiaye, Chen, Baoyou + 8 more
Transformers rely on explicit positional encoding to model structure in data. While Rotary Position Embedding (RoPE) excels in 1D domains, its application to image generation reveals significant limitations such as fine-grained spatial relation modeling, color cues, and object counting. This paper identifies key…
Thijs L van der Plas, Jacob JW Bakermans, Vishal Nedungadi, Gabrielė Tijūnaitytė + 2 more
Earth embedding models transform Earth observation data into embeddings uniquely tied to locations on the Earth's surface. These models are typically evaluated in isolation, comparing the downstream task performance across different Earth embeddings. However, spatially aligned embeddings can naturally be fused…
Doaa Sami Khafaga, El-Sayed M. El-kenawy, Nima Khodadadi, Marwa M. Eid + 1 more
Image watermarking is an important extension of intellectual property protection that facilitates the identification and authentication of multimedia content. This paper aims to improve and optimize image watermarking techniques to ensure effective image protection regardless of image size or format. The proposed…
Beiji Lu
Synonymous codons encode the same amino acid yet are used non-randomly across genomes, a phenomenon with well-documented functional consequences for translation efficiency and mRNA stability. Whether the information embedded in synonymous codon choice is recoverable from the internal representations of in-dependently…
Authors not listed
Chemical data is fundamentally sparse, with molecular structures serving as database keys for countless properties. Current machine learning methods map structures to properties with remarkable accuracy, yet they do not leverage available property information when predicting unknowns, creating unutilized partial…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Authors not listed
Sequence-defined oligomers offer programmable molecular architectures with potential in data storage, authentication, and anticounterfeiting. However, their deployment in real-world materials has been constrained by their low scale, limited thermal resilience and the need for specialized analytical methods. Here we…