24 papers · ranked by Valyu relevance
Fabio Cumbo, Kabir Dhillon, Jayadev Joshi, Davide Chicco + 2 more
Viral species classification is crucial for understanding viral evolution, epidemiology, and developing effective diagnostics and treatments. Traditional methods often rely on sequence similarity, which can be challenging for rapidly evolving viruses. Pangenomes, offering a comprehensive representation of species’…
Anders Høst-Madsen, Jun Zhang
—This paper has dual aims. First is to develop practical universal coding methods for unlabeled graphs. Second is to use these for graph anomaly detection. The paper develops two coding methods for unlabeled graphs: one based on the degree distribution, the second based on the triangle distribution. It is shown that…
Muhammad Umair, Young-Koo Lee
Graph data are pervasive worldwide, e.g., social networks, citation networks, and web graphs. A real-world graph can be huge and requires heavy computational and storage resources for processing. Various graph compression techniques have been presented to accelerate the processing time and utilize memory efficiently.…
Hung T. Nguyen, Pierre Jinghong Liang, Leman Akoglu
Within a large database G containing graphs with labeled nodes and directed, multi-edges; how can we detect the anomalous graphs? Most existing work are designed for plain (unlabeled) and/or simple (unweighted) graphs. We introduce CODEtect, the first approach that addresses the anomaly detection task for graph…
Mojtaba Abolfazli, Anders Høst-Madsen, Jun Zhang, András Bratincsák
—Many multivariate data such as social and biological data exhibit complex dependencies that are best characterized by graphs. Unlike sequential data, graphs are, in general, unordered structures. This means we can no longer use classic, sequential-based compression methods on these graph-based data. Therefore, it is…
Giorgos Bouritsas, Andreas Loukas, Nikolaos Karalias, Michael M. Bronstein
'Michael M. Bronstein'] Can we use machine learning to compress graph data? The absence of ordering in graphs poses a significant challenge to conventional compression algorithms, limiting their attainable gains as well as their ability to discover relevant patterns. On the other hand, most graph compression approaches…
Prathyush Poduval, Haleh Alimohamadi, Ali Zakeri, Farhad Imani + 3 more
'M. Hassan Najafi' 'Tony Givargis' 'Mohsen Imani'] Memorization is an essential functionality that enables today's machine learning algorithms to provide a high quality of learning and reasoning for each prediction. Memorization gives algorithms prior knowledge to keep the context and define confidence for their…
Luca Cappelletti, Tommaso Fontana, Elena Casiraghi, Vida Ravanmehr + 7 more
'Tiffany J. Callahan' 'Carlos Cano' 'Marcin P. Joachimiak' 'Christopher J. Mungall' 'Peter N. Robinson' 'Justin Reese' 'Giorgio Valentini'] Graph representation learning methods opened new avenues for addressing complex, real-world problems represented by graphs. However, many graphs used in these applications comprise…
Alexis Bénichou, Jean-Baptiste Masson, Christian L. Vestergaard, Fabrizio De Vico Fallani
Physical and functional constraints on biological networks lead to complex topological patterns across multiple scales in their organization. A particular type of higher-order network feature that has received considerable interest is network motifs, defined as statistically regular subgraphs. These may implement…
Lloyd Allison
TR #2014/2771 . This report concerns the information content of a graph, optionally conditional on one or more background, "common knowledge" graphs. It describes an algorithm to estimate this information content, and includes some examples based on chemical compounds. keywords: Graph, network, complexity, information…
Md Toki Tahmid, Tanjeem Azwad Zaman, Mohammad Saifur Rahman
Understanding complex graph-structured data is a cornerstone of modern research in fields like cheminformatics and bioinformatics, where molecules and biological systems are naturally represented as graphs. However, traditional graph neural networks (GNNs) often fall short by focusing mainly on node features while…
Jordan M. Eizenga, Adam M. Novak, Emily Kobayashi, Flavia Villani + 6 more
Pangenomics is a growing field within computational genomics. Many pangenomic analyses use bidirected sequence graphs as their core data model. However, implementing and correctly using this data model can be difficult, and the scale of pangenomic data sets can be challenging to work at. These challenges have impeded…
Van Thuy Hoang, Hyeon-Ju Jeon, Eun-Soon You, Yoewon Yoon + 3 more
Graphs are data structures that effectively represent relational data in the real world. Graph representation learning is a significant task since it could facilitate various downstream tasks, such as node classification, link prediction, etc. Graph representation learning aims to map graph entities to low-dimensional…
Peter Heringer, Daniel Doerr
Pangenome graphs offer a compact and comprehensive representation of genomic diversity, improving tasks such as variant calling, genotyping, and other downstream analyses. Although the underlying graph structures scale sublinearly with the number of haplotypes, the widely used GFA file format suffers from rapidly…
Zhiyuan Ding, Alex Baras
Recent advances in computation pathology have seen the development of various forms of foundational models that have enabled high-quality, generalpurpose feature extraction from tissue patches. However, most of these models are somewhat limited in their ability to capture cell-to-cell spatial relationships essential…
Xueyuan Chen, Shangzhe Li, Yanchun Liang
Due to the success observed in deep neural networks with contrastive learning, there has been a notable surge in research interest in graph contrastive learning, primarily attributed to its superior performance in graphs with limited labeled data. Within contrastive learning, the selection of a “view” dictates the…
Vladimir Kondratyev, Marian Dryzhakov, Timur Gimadiev, Dmitriy Slutskiy
In this work, we provide further development of the junction tree variational autoencoder (JT VAE) architecture in terms of implementation and application of the internal feature space of the model. Pretraining of JT VAE on a large dataset and further optimization with a regression model led to a latent space that can…
Authors not listed
Computational methods for predictive modeling have been increasingly utilized in the early stages of drug discovery to supplement high-throughput screening. The advent of highly efficient and complex machine learning architectures necessitates new methods of collating the plethora of topological, geometrical, and…
David Buterez, Jon Paul Janet, Steven Kiddle, Pietro Liò
We investigate the potential of graph neural networks for transfer learning and improving molecular property prediction on sparse and expensive to acquire high-fidelity data by leveraging low-fidelity measurements as an inexpensive proxy for a targeted property ofinterest. This problem arises in discovery processes…
Cailum Stienstra, Liam Hebert, Patrick Thomas, Alexander Haack + 2 more
Given that Infrared (IR) spectroscopy is a crucial tool in various chemical and forensic domains, improved in silico methods for predicting experimental spectra are needed due to the time and accuracy limitations of ab initio methods. We employ Graphormer, a graph neural network (GNN) transformer, to predict IR spectra…
Ian T. Hoffecker, Yunshi Yang, Giulio Bernardinelli, Pekka Orponen + 1 more
Barcoded DNA polony amplification techniques provide a means to impart a unique sequence identity onto specific locations of a surface wafer or chip. We describe a method whereby micro-scale spatial information such as the relative positions of biomolecules on a surface can be transferred to a sequence-based format and…
Maria Boulougouri, Pierre Vandergheynst, Daniel Probst
Computational representation of molecules can take many forms, including graphs, stringencodings of graphs, binary vectors, or learned embeddings in the form of real-valued vectors. These representations are then used in downstream classification and regression tasks using a wide range of machine-learning models.…
Authors not listed
A directed graph (or digraph) consists of a finite vertex set 𝑉 and a set of ordered edges 𝐸 ⊆ 𝑉 × 𝑉, each edge (𝑢, 𝑣) indicating a one-way connection from 𝑢 (source) to 𝑣 (target). A bidirected graph is a generalization of an undirected graph where each edge is assigned a direction at each of its endpoints…
Authors not listed
Graph Neural Networks (GNNs) have emerged as a powerful tool in predicting molecular properties based on structural data. While GNNs excel in identifying local patterns within molecules, their ability to capture global properties remains limited due to inherent structural challenges such as oversmoothing and their…