24 papers · ranked by Valyu relevance
Muhammad Umair, Young-Koo Lee
Graph data are pervasive worldwide, e.g., social networks, citation networks, and web graphs. A real-world graph can be huge and requires heavy computational and storage resources for processing. Various graph compression techniques have been presented to accelerate the processing time and utilize memory efficiently.…
Kenta Yanagiya, Junya Hara, Hiroshi Higashi, Yuichi Tanaka + 1 more
In this paper, we propose a compression framework for weighted graphs in which the graph topology is transmitted losslessly and edge weights are compressed lossily. A challenge in the lossy compression of edge weights is that the underlying relationships between edges are ambiguous. To address this issue, we first…
Youngchun Kwon, Dongseon Lee, Youn-Suk Choi, Kyoham Shin + 1 more
'Seokho Kang'] Recently, deep learning has been successfully applied to molecular graph generation. Nevertheless, mitigating the computational complexity, which increases with the number of nodes in a graph, has been a major challenge. This has hindered the application of deep learning-based molecular graph generation…
Muhammad Ifte Islam, Farhan Tanvir, Ginger Johnson, Esra Akbas + 1 more
Network embedding that encodes structural information of graphs into a low-dimensional vector space has been proven to be essential for network analysis applications, including node classification and community detection. Although recent methods show promising performance for various applications, graph embedding still…
Ingo Schilken, Harun Mustafa, Gunnar Rätsch, Carsten Eickhoff + 1 more
Technological advancements in high throughput DNA sequencing have led to an exponential growth of sequencing data being produced and stored as a byproduct of biomedical research. Despite its public availability, a majority of this data remains inaccessible to the research community through a lack efficient data…
Morihiro Hayashida, Tatsuya Akutsu
Background Comparison of various kinds of biological data is one of the main problems in bioinformatics and systems biology. Data compression methods have been applied to comparison of large sequence data and protein structure data. Since it is still difficult to compare global structures of large biological networks…
Luca Versari, Iulia M. Comşa, Alessio Conte, Roberto Grossi
Zuckerli is a scalable compression system meant for large real-world graphs. Graphs are notoriously challenging structures to store efficiently due to their linked nature, which makes it hard to separate them into smaller, compact components. Therefore, effective compression is crucial when dealing with large graphs…
Muhammad Irfan Yousuf, Izza Anwer, Muhammad Zeeshan Abid
Real-world graphs are massive in size and we need a huge amount of space to store them. Graph compression allows us to compress a graph so that we need a lesser number of bits per link to store it. Of many techniques to compress a graph, a typical approach is to find clique-like caveman or traditional communities in a…
Tangina Sultana, Young-Koo Lee, Wookey Lee
The explosive volume of semantic data published in the Resource Description Framework (RDF) data model demands efficient management and compression with better compression ratio and runtime. Although extensive work has been carried out for compressing the RDF datasets, they do not perform well in all dimensions.…
Ferhat Ay, Michael Dang, Tamer Kahveci
Metabolic network alignment is a system scale comparative analysis that discovers important similarities and differences across different metabolisms and organisms. Although the problem of aligning metabolic networks has been considered in the past, the computational complexity of the existing solutions has so far…
Maciej Besta, Torsten Hoefler
Various graphs such as web or social networks may contain up to trillions of edges. Compressing such datasets can accelerate graph processing by reducing the amount of I/O accesses and the pressure on the memory subsystem. Yet, selecting a proper compression method is challenging as there exist a plethora of…
Yun Chang, Luca Ballotta, Luca Carlone
—For a multi-robot team that collaboratively explores an unknown environment, it is of vital importance that the collected information is efficiently shared among robots in order to support exploration and navigation tasks. Practical constraints of wireless channels, such as limited bandwidth, urge robots to carefully…
Sebastian Maneth, Fabian Peternek
We present an informal survey (meant to accompany another paper) on graph compression methods. We focus on lossless methods, briefly list available approaches, and compare them where possible or give some indicators on their compression ratios. We also mention some relevant results from the field of lossy compression…
Robin Lamarche-Perrin
Graph compression is a data analysis technique that consists in the replacement of parts of a graph by more general structural patterns in order to reduce its description length. It notably provides interesting exploration tools for the study of real, large-scale, and complex graphs which cannot be grasped at first…
Daniel Danciu, Mikhail Karasikov, Harun Mustafa, André Kahles + 1 more
Since the amount of published biological sequencing data is growing exponentially, efficient methods for storing and indexing this data are more needed than ever to truly benefit from this invaluable resource for biomedical research. Labeled de Bruijn graphs are a frequently-used approach for representing large sets of…
Jouni Sirén, Benedict Paten
Pangenome graphs representing aligned genome assemblies are being shared in the text-based Graphical Fragment Assembly format. As the number of assemblies grows, there is a need for a file format that can store the highly repetitive data space-efficiently. We propose the GBZ file format based on data structures used in…
Yutong Qiu, Carl Kingsford
The size of a genome graph — the space required to store the nodes, their labels and edges — affects the efficiency of operations performed on it. For example, the time complexity to align a sequence to a graph without a graph index depends on the total number of characters in the node labels and the number of edges in…
Nguyen Gia Bach, Chanh Minh Tran, Tho Nguyen Duc, Phan Xuan Tan + 2 more
'Eiji Kamioka' 'Yitzhak Yitzhaky'] In light field compression, graph-based coding is powerful to exploit signal redundancy along irregular shapes and obtains good energy compaction. However, apart from high time complexity to process high dimensional graphs, their graph construction method is highly sensitive to the…
Mikhail Karasikov, Harun Mustafa, Amir Joudaki, Sara Javadzadeh No + 2 more
High-throughput DNA sequencing data is accumulating in public repositories, and efficient approaches for storing and indexing such data are in high demand. In recent research, several graph data structures have been proposed to represent large sets of sequencing data and allow for efficient query of sequences. In…
Jan Weinreich, Daniel Probst
In recent years, natural language processing approaches to machine learning, most prominently deep neural network-based transformers, have been extensively applied to molecular classification and regression tasks, including the prediction of pharmacokinetic and quantum-chemical properties. However, models based on deep…
Authors not listed
Real-world datasets in chemical engineering and bioengineering processes--such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials--can often be unlabelled or disorganized, rendering the training of existing supervised learning models ineffective at learning the…
Authors not listed
Genetic Algorithms are a powerful method to solve optimization problems with complex cost functions over vast search spaces that rely in particular on recombining parts of previous solutions. Crossover operators play a crucial role in this context. Here, we describe a large class of these operators designed for…
Dong Hyeon Mok, Jongseung Kim, Seoin Back
To realize renewable and sustainable energy cycle, there has been a lot of effort put into discovering catalysts with desired properties from a large chemical space. To achieve this goal, several screening strategies have been proposed, most of which require validation of thermodynamic stability and synthesizability of…
Authors not listed
Rapid and robust simulation of chemical processes is critical to conduct process design, optimization, techno-economic analysis, and sustainability analysis. Yet, efficiently solving simulation models remains a challenge due to the highly coupled and nonlinear nature of the underlying algebraic equations that capture…