24 papers · ranked by Valyu relevance
Muhammad Umair, Young-Koo Lee
Graph data are pervasive worldwide, e.g., social networks, citation networks, and web graphs. A real-world graph can be huge and requires heavy computational and storage resources for processing. Various graph compression techniques have been presented to accelerate the processing time and utilize memory efficiently.…
Kenta Yanagiya, Junya Hara, Hiroshi Higashi, Yuichi Tanaka + 1 more
In this paper, we propose a compression framework for weighted graphs in which the graph topology is transmitted losslessly and edge weights are compressed lossily. A challenge in the lossy compression of edge weights is that the underlying relationships between edges are ambiguous. To address this issue, we first…
Maximilien Danisch, Ioannis Panagiotas, Lionel Tabourier
In order to manage massive graphs in practice, it is often necessary to resort to graph compression, which aims at reducing the memory used when storing and processing the graph. Efficient compression methods have been proposed in the literature, especially for web graphs. In most cases, they are combined with a vertex…
Yun Chang, Luca Ballotta, Luca Carlone
—For a multi-robot team that collaboratively explores an unknown environment, it is of vital importance that the collected information is efficiently shared among robots in order to support exploration and navigation tasks. Practical constraints of wireless channels, such as limited bandwidth, urge robots to carefully…
Yufeng Zhang, Weiyao Lin, Wenrui Dai, Huabin Liu + 1 more
The scene graph is a new data structure describing objects and their pairwise relationship within image scenes. As the size of scene graph in vision applications grows, how to losslessly and efficiently store such data on disks or transmit over the network becomes an inevitable problem. However, the compression of…
Amirmohammad Farzaneh, Justin P. Coon, Mihai-Alin Badiu, Narsis A. Kiani + 2 more
'Narsis A. Kiani' 'Hector Zenil' 'Jesper Tegnér'] Throughout the years, measuring the complexity of networks and graphs has been of great interest to scientists. The Kolmogorov complexity is known as one of the most important tools to measure the complexity of an object. We formalized a method to calculate an upper…
Akshar Chavan, Sanaz Rabina, Daniel Grosu, Marco Brocanelli
Reducing the running time of graph algorithms is vital for tackling real-world problems such as shortest paths and matching in large-scale graphs, where path information plays a crucial role. This paper addresses this critical challenge of reducing the running time of graph algorithms by proposing a new graph…
Kenta Yanagiya, Junya Hara, Hiroshi Higashi, Yuichi Tanaka + 1 more
'Antonio Ortega'] This paper proposes a compression framework for adjacency matrices of weighted graphs based on graph filter banks. Adjacency matrices are widely used mathematical representations of graphs and are used in various applications in signal processing, machine learning, and data mining. In many problems of…
Wenfei Fan, Yuanhao Li, Muyang Liu, Can Lu
This paper proposes a scheme to reduce big graphs to small graphs. It contracts obsolete parts and regular structures into supernodes. The supernodes carry a synopsis $S_\mathcal{Q}$ for each query class $\mathcal{Q}$ in use, to abstract key features of the contracted parts for answering queries of $\mathcal{Q}$.…
Tangina Sultana, Young-Koo Lee, Wookey Lee
The explosive volume of semantic data published in the Resource Description Framework (RDF) data model demands efficient management and compression with better compression ratio and runtime. Although extensive work has been carried out for compressing the RDF datasets, they do not perform well in all dimensions.…
Timothé Rouzé, Rayan Chikhi, Antoine Limasset
Petabases of sequencing data in the Sequence Read Archive (SRA) present a significant challenge for holistic reanalysis due to their sheer volume. Recent efforts have assembled this data into terabytes of unitigs, an efficient k-mer set representation that can reduce data size by an order of magnitude. However, these…
Drew DeHaas, Ziqing Pan, Xinzhu Wei
Computational analysis of a large number of genomes requires a data structure that can represent the dataset compactly while also enabling efficient operations on variants and samples. Current practice is to store large-scale genetic polymorphism data using tabular data structures and file formats, where rows and…
Jouni Sirén, Benedict Paten
Pangenome graphs representing aligned genome assemblies are being shared in the text-based Graphical Fragment Assembly format. As the number of assemblies grows, there is a need for a file format that can store the highly repetitive data space-efficiently. We propose the GBZ file format based on data structures used in…
Nguyen Gia Bach, Chanh Minh Tran, Tho Nguyen Duc, Phan Xuan Tan + 2 more
'Eiji Kamioka' 'Yitzhak Yitzhaky'] In light field compression, graph-based coding is powerful to exploit signal redundancy along irregular shapes and obtains good energy compaction. However, apart from high time complexity to process high dimensional graphs, their graph construction method is highly sensitive to the…
Dimitris Floros, Nikos Pitsianis, Xiaobai Sun
Locality and Graph Compression Authors: Dimitris Floros, Nikos Pitsianis, Xiaobai Sun Title: Algebraic Vertex Ordering of a Sparse Graph for Adjacency Access Locality and Graph Compression Authors: Dimitris Floros, Nikos Pitsianis, Xiaobai Sun Content: ## I. INTRODUCTION In a modern data, knowledge, or information…
Jouni Sirén, Benedict Paten
Existing pangenome file formats are designed for batch processing. Graphs must be loaded into memory, and alignment files must be read sequentially. Indexed file formats that can be used directly from disk would be more appropriate for interactive applications. We propose GBZ-base and GAF-base — SQLite-backed file…
Phanindra Reddy Madduru, Bijo Thomas
This paper proposes a preprocessing framework for optimizing large-scale graph database ingestion through intelligent edge filtering based on value ranking. We combine adapted PageRank algorithms with business-specific metrics and edge type importance to evaluate and rank edges, enabling selective retention of…
Amatur Rahman, Yoann Dufresne, Paul Medvedev
A colored de Bruijn graph (also called a set of k-mer sets), is a set of k-mers with every k-mer assigned a set of colors. Colored de Bruijn graphs are used in a variety of applications, including variant calling, genome assembly, and database search. However, their size has posed a scalability challenge to algorithm…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…
Authors not listed
Real-world datasets in chemical engineering and bioengineering processes--such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials--can often be unlabelled or disorganized, rendering the training of existing supervised learning models ineffective at learning the…
Authors not listed
Genetic Algorithms are a powerful method to solve optimization problems with complex cost functions over vast search spaces that rely in particular on recombining parts of previous solutions. Crossover operators play a crucial role in this context. Here, we describe a large class of these operators designed for…
Dong Hyeon Mok, Jongseung Kim, Seoin Back
To realize renewable and sustainable energy cycle, there has been a lot of effort put into discovering catalysts with desired properties from a large chemical space. To achieve this goal, several screening strategies have been proposed, most of which require validation of thermodynamic stability and synthesizability of…
Authors not listed
Rapid and robust simulation of chemical processes is critical to conduct process design, optimization, techno-economic analysis, and sustainability analysis. Yet, efficiently solving simulation models remains a challenge due to the highly coupled and nonlinear nature of the underlying algebraic equations that capture…
Runzhao Yang, Tingxiong Xiao, Yuxiao Cheng, Anan Li + 7 more
Efficient storage and sharing of massive biomedical data would open up their wide accessibility to different institutions and disciplines. However, compressors tailored for natural photos/videos are rapidly limited for biomedical data, while emerging deep learning based methods demand huge training data and are…