24 papers · ranked by Valyu relevance
Mattia Cervellini, Blerina Sinaimeri, Catherine Matias, Alessio Martino
Metabolic networks are complex systems that describe the biochemical reactions within an organism through pairwise interactions between chemical compounds. While this representation is widely used to study biological function, it fails to capture the full structure of metabolic reactions, many of which involve more…
Mattia Cervellini, Blerina Sinaimeri, Catherine Matias, Alessio Martino
Metabolic networks are complex systems that describe the bio-chemical reactions within an organism through pairwise interactions between chemical compounds. While this representation is widely used to study biolog-ical function, it fails to capture the full structure of metabolic reactions, many of which involve more…
Jeffrey Zhong, Lechuan Li, Ruth Dannenfelser, Vicky Yao
Gene embeddings have emerged as transformative tools in computational biology, enabling the efficient translation of complex biological datasets into compact vector representations. This study presents a comprehensive benchmark by evaluating 38 classic and state-of-the-art gene embedding methods across a spectrum of…
Maximilian Nickel, Douwe Kiela
Representation learning has become an invaluable approach for learning from symbolic data such as text and graphs. However, while complex symbolic datasets often exhibit a latent hierarchical structure, state-of-the-art methods typically learn embeddings in Euclidean vector spaces, which do not account for this…
Ilya Makarov, Dmitrii Kiselev, Nikita Nikitinsky, Lovro Subelj + 1 more
Dealing with relational data always required significant computational resources, domain expertise and task-dependent feature engineering to incorporate structural information into a predictive model. Nowadays, a family of automated graph feature engineering techniques has been proposed in different streams of…
Miroslav Kratochvíl, Abhishek Koladiya, Jana Balounova, Vendula Novosadova + 4 more
Efficient unbiased data analysis is a major challenge for laboratories handling large flow and mass cytometry datasets. We present EmbedSOM, a non-linear embedding algorithm based on FlowSOM that improves the analysis by providing high-performance embedding method for the cytometry data. The algorithm is designed for…
Miroslav Kratochvíl, Abhishek Koladiya, Jiří Vondrášek
EmbedSOM is a simple and fast dimensionality reduction algorithm, originally developed for its applications in single-cell cytometry data analysis. We present an updated version of EmbedSOM, viewed as an algorithm for landmark-directed embedding enrichment, and demonstrate that it works well even with manifold-learning…
Nada Lavrač, Blaž Škrlj, Marko Robnik-Šikonja
Data preprocessing is an important component of machine learning pipelines, which requires ample time and resources. An integral part of preprocessing is data transformation into the format required by a given learning algorithm. This paper outlines some of the modern data processing techniques used in relational…
Jeffrey Zhong, Lechuan Li, Ruth Dannenfelser, Vicky Yao
Accurate, data-driven representations of genes are critical for interpreting high-throughput biological data, yet no consensus exists on the most effective embedding strategy for common functional prediction tasks. Here, we present a systematic comparison of 38 gene embedding methods derived from amino acid sequences…
Zhiheng Huang, Davis Liang, Peng Xu, Bing Xiang
Transformer architectures rely on explicit position encodings in order to preserve a notion of word order. In this paper, we argue that existing work does not fully utilize position information. For example, the initial proposal of a sinusoid embedding is fixed and not learnable. In this paper, we first review absolute…
Ling Cai, Krzysztof Janowicz, Rui Zhu, Gengchen Mai + 2 more
Qualitative spatial/temporal reasoning (QSR/QTR) plays a key role in research on human cognition, e.g., as it relates to navigation, as well as in work on robotics and artificial intelligence. Although previous work has mainly focused on various spatial and temporal calculi, more recently representation learning…
Paola Lecca, Michela Lecca
Graphs are used as a model of complex relationships among data in biological science since the advent of systems biology in the early 2000. In particular, graph data analysis and graph data mining play an important role in biology interaction networks, where recent techniques of artificial intelligence, usually…
Gianluca Tirimbo, Vivek Sundaram, Björn Baumeier
Many-body Green's function theory in the GW approximation with the Bethe--Salpeter equation (BSE) provides a powerful framework for the first-principles calculations of single-particle and electron-hole excitations in perfect crystals and molecules alike. Application to complex molecular systems, e.g., solvated dyes…
Simchoni, Giora, Rosset, Saharon
We present MMbeddings, a probabilistic embedding approach that reinterprets categorical embeddings through the lens of nonlinear mixed models, effectively bridging classical statistical theory with modern deep learning. By treating embeddings as latent random effects within a variational autoencoder framework, our…
Authors not listed
High-level quantum mechanical (QM) simulations provide accurate electronic information of chemical systems but scale unfavourably with system size, making calculations of applied systems challenging. Hierarchical quantum mechanics in quantum mechanics embedding (QM/QM) addresses this issue by localising the highly…
Ujjal Kr Dutta, Mehrtash Harandi, Chandra Sekhar Chellu
For challenging machine learning problems such as zero-shot learning and fine-grained categorization, embedding learning is the machinery of choice because of its ability to learn generic notions of similarity, as opposed to class-specific concepts in standard classification models. Embedding learning aims at learning…
Xiaoli Huang, Haibo Chen, Zheng Zhang, Donald J. Jacobs + 1 more
'Sotiris Kotsiantis'] Hash is one of the most widely used methods for computing efficiency and storage efficiency. With the development of deep learning, the deep hash method shows more advantages than traditional methods. This paper proposes a method to convert entities with attribute information into embedded vectors…
Authors not listed
Hybrid machine-learning/molecular-mechanics (ML/MM) methods extend the classical QM/MM paradigm by replacing the quantum desription with neural network interatomic potentials trained to reproduce accurately quantum-mechanical (QM) results. By describing only the chemically active region with ML and the surrounding…
Shiwei Li, Huifeng Guo, Xing Tang, Ruiming Tang + 3 more
'Ruixuan Li' 'Rui Zhang'] To alleviate the problem of information explosion, recommender systems are widely deployed to provide personalized information filtering services. Usually, embedding tables are employed in recommender systems to transform high-dimensional sparse one-hot vectors into dense real-valued…
Kevin Z. Lin, Jing Lei, Kathryn Roeder
Scientists often embed cells into a lower-dimensional space when studying single-cell RNA-seq data for improved downstream analyses such as developmental trajectory analyses, but the statistical properties of such non-linear embedding methods are often not well understood. In this article, we develop the eSVD…
Christoph Jacob, Johannes Neugebauer
The past years since the publication of our review on subsystem density-functional theory (sDFT) [WIREs Comput. Mol. Sci. 2014, 4:325--362] have witnessed a rapid development and diversification of quantum mechanical fragmentation and embedding approaches related to sDFT and frozen-density embedding (FDE). In this…
Kan Wu, Houwen Peng, Minghao Chen, Jianlong Fu + 1 more
Relative position encoding (RPE) is important for transformer to capture sequence ordering of input tokens. General efficacy has been proven in natural language processing. However, in computer vision, its efficacy is not well studied and even remains controversial, e.g., whether relative position encoding can work…
Authors not listed
Real-world datasets in chemical engineering and bioengineering processes--such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials--can often be unlabelled or disorganized, rendering the training of existing supervised learning models ineffective at learning the…
Authors not listed
Elucidating Collective Variables (CVs) for biomolecular dynamics is crucial for understanding numerous biological processes. By leveraging the tensor-train data structure, a multilinear version of the AMUSE (Algorithm for Multiple Unknown Signals) algorithm for Koopman approximation (AMUSEt) was recently developed to…