26 papers · ranked by Valyu relevance
Dongqi Fu, Jingrui He
Graph structures have attracted much research attention for carrying complex relational information. Based on graphs, many algorithms and tools are proposed and developed for dealing with real-world tasks such as recommendation, fraud detection, molecule design, etc. In this paper, we first discuss three topics of…
Matthias van der Hallen, S. V. Paramonov, Michaël Leuschel, Gerda Janssens
'Gerda Janssens'] Abstract. Many problems, especially those with a composite structure, can naturally be expressed in higher order logic. From a KR perspective modeling these problems in an intuitive way is a challenging task. In this paper we study the graph mining problem as an example of a higher order problem. In…
Cheng Zhao, Zhibin Zhang, Peng Xu, Tianqi Zheng + 1 more
Graph mining is one of the most important categories of graph algorithms. However, exploring the subgraphs of an input graph produces a huge amount of intermediate data. The "think like a vertex" programming paradigm, pioneered by Pregel, cannot readily formulate mining problems, which is designed to produce graph…
Aida Mrzic, Pieter Meysman, Wout Bittremieux, Pieter Moris + 3 more
Searching for interesting common subgraphs in graph data is a well-studied problem in data mining. Subgraph mining techniques focus on the discovery of patterns in graphs that exhibit a specific network structure that is deemed interesting within these data sets. The definition of which subgraphs are interesting and…
Sebastian Keller, Pauli Miettinen, Olga V. Kalinina
Identification of biologically relevant motifs in proteins is a long-standing problem in bioinformatics, especially when considering distantly related proteins where sequence analysis alone becomes increasingly difficult. Here we present a novel approach to identify such motifs in protein three-dimensional structures…
Kasra Jamshidi, Rakesh Mahadasa, Keval Vora
Graph mining workloads aim to extract structural properties of a graph by exploring its subgraph structures. General purpose graph mining systems provide a generic runtime to explore subgraph structures of interest with the help of userdefined functions that guide the overall exploration process. However, the…
Belgin Ergenç Bostanoğlu, Nourhan Abuzayed, Bilal Alatas
Frequent subgraph mining (FSM) is an essential and challenging graph mining task used in several applications of the modern data science. Some of the FSM algorithms have the objective of finding all frequent subgraphs whereas some of the algorithms focus on discovering frequent subgraphs approximately. On the other…
Aaron J. Gutknecht, Michael Wibral
We describe how the recently introduced method of significant subgraph mining can be employed as a useful tool in network comparison. It is applicable whenever the goal is to compare two sets of unweighted graphs and to determine differences in the processes that generate them. We provide an extension of the method to…
Ali Jazayeri, Christopher C. Yang
Motifs are the fundamental components of complex systems. The topological structure of networks representing complex systems and the frequency and distribution of motifs in these networks are intertwined. The complexities associated with graph and subgraph isomorphism problems, as the core of frequent subgraph mining…
Carlos H. C. Teixeira, Alexandre J. Fonseca, Marco Serafini, Georgos Siganos + 2 more
Distributed data processing platforms such as MapReduce and Pregel have substantially simplified the design and deployment of certain classes of distributed graph analytics algorithms. However, these platforms do not represent a good match for distributed graph mining problems, as for example finding frequent subgraphs…
Yike Liu, Abhilash Dighe, Tara Safavi, Danai Koutra
While advances in computing resources have made processing enormous amounts of data possible, human ability to identify patterns in such data has not scaled accordingly. Efficient computational methods for condensing and simplifying data are thus becoming vital for extracting actionable insights. In particular, while…
Karthikeyan Rajendran, Assimakis Kattis, Alexander Holiday, Risi Kondor + 1 more
'Risi Kondor' 'Ioannis G. Kevrekidis'] Abstract We discuss the problem of extending data mining approaches to cases in which data points arise in the form of individual graphs. Being able to find the intrinsic low-dimensionality in ensembles of graphs can be useful in a variety of modeling contexts, especially when…
Charles Packer, Lawrence B. Holder
A massive amount of data generated today on platforms such as social networks, telecommunication networks, and the internet in general can be represented as graph streams. Activity in a network's underlying graph generates a sequence of edges in the form of a stream; for example, a social network may generate a graph…
Mohammed Alokshiya, Saeed Salem, Fidaa Abed
Background Real biological and social data is increasingly being represented as graphs. Pattern-mining-based graph learning and analysis techniques report meaningful biological subnetworks that elucidate important interactions among entities. At the backbone of these algorithms is the enumeration of pattern space.…
Daniel Walke, Daniel Micheel, Kay Schallert, Thilo Muth + 3 more
'David Broneske' 'Gunter Saake' 'Robert Heyer'] Title: Abstract The increasing amount and complexity of clinical data require an appropriate way of storing and analyzing those data. Traditional approaches use a tabular structure (relational databases) for storing data and thereby complicate storing and retrieving…
Alokkumar Jha, Yasar Khan, Ratnesh Sahay
Prediction of metastatic sites from the primary site of origin is a impugn task in breast cancer (BRCA). Multi-dimensionality of such metastatic sites - bone, lung, kidney, and brain, using large-scale multi-dimensional Poly-Omics (Transcriptomics, Proteomics and Metabolomics) data of various type, for example, CNV…
Tom C. Freeman, Sebastian Horsewell, Anirudh Patir, Josh Harling-Lee + 5 more
Quantitative and qualitative data derived from the analysis of genomes, genes, proteins or metabolites from tissue or cells are currently generated in huge volumes during biomedical research. Graphia is an open-source platform created for the graph-based analysis of such complex data, e.g. transcriptomics, proteomics…
M. Zanin, D. Papo, P. A. Sousa, E. Menasalvas + 3 more
The increasing power of computer technology does not dispense with the need to extract meaningful in-formation out of data sets of ever growing size, and indeed typically exacerbates the complexity of this task. To tackle this general problem, two methods have emerged, at chronologically different times, that are now…
Authors not listed
Background: Chemical reactions form intricate, highly connected networks whose exploration is essential for discovering more efficient and sustainable synthetic routes. As reaction data from literature, patents, and high‑throughput experimentation continue to surge, so does the need for tools that can collectively…
Gergely Zahoránszky-Kőhalmi, Brandon Walker, Nathan Miller, Brett Yang + 11 more
The recent SmartGraph platform facilitates the execution of complex drug-discovery workflows with ease in the network-pharmacology paradigm. However, at the time of its publication, we identified the need for the development of an Application Programming Interface (API) that could promote biomedical data integration…
Jana Weber, Pietro Lio’, Alexei Lapkin
Networks of chemical reactions represent relationships between molecules within chemical supply chains and promise to enhance planning of multi-step synthesis routes from bio-renewable feedstocks. This study aims to identify strategic molecules in chemical reaction networks that may potentially play a significant role…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Alexander Smith, Spencer Runde, Alex Chew, Atharva Kelkar + 3 more
Molecular dynamics (MD) simulations are used in diverse scientific and engineering fields such as drug discovery, materials design, separations, biological systems, and reaction engineering. These simulations generate highly complex datasets that capture the 3D spatial positions, dynamics, and interactions of thousands…
Lionel Zoubritzky, François-Xavier Coudert
We present here an open-source Julia library for the topological identification of crystalline materials, with algorithmic and computational improvements over the previously available software in the field, resulting in a speed increase of one order of magnitude. This new algorithm and implementation can therefore be…
Authors not listed
We present a unified, set–theoretic framework that extends molecular graphs to hypergraphs and superhypergraphs via iterated power sets. We define Molecular Graphs, Molecular HyperGraphs, and Molecular SuperHyperGraphs, and develop four complements over them: Weighted, Rough, Neural, and Multipolar frameworks. We prove…
Pulan Yu
Associative classification mining (ACM) integrating association rule mining and classification has become a significant tool for knowledge discovery, especially in the chemical domain. Its major advantage is providing high accuracy as well as chemically interpretable models. Additionally, it is able to find…