14 papers · ranked by Valyu relevance
James P. Bagrow, Yong‐Yeol Ahn
The deluge of network datasets demands a standard way to effectively and succinctly summarize network datasets. Building on similar efforts to standardize the documentation of models and datasets in machine learning, here we propose network cards, short summaries of network datasets that can capture not only the basic…
Tatjana Welzer, Johann Eder, Vili Podgorelec, Robert Wrembel + 6 more
'Mirjana Ivanonvic' 'Johann Gamper' 'Mikołaj Morzy' 'Theodoros Tzouramanis' 'Jérôme Darmont' 'Aida Kamišalić Latifić'] Abstract. Over the past decade, the data lake concept has emerged as an alternative to data warehouses for storing and analyzing big data. A data lake allows storing data without any predefined schema.…
Yahya Sa'd, Renzo Angles, Vojtech Merunka, Roberto Garcia + 2 more
Property-graph schemas often contain descriptive properties that recur across heterogeneous nodes and edges, yet schema designers lack a clear method for deciding whether such properties should remain embedded or be treated as reusable metadata structures. This paper addresses this design-stage problem within a…
Sepehr Sadoughi, Nikolay Yakovets, George Fletcher
and Reification Authors: ['Sepehr Sadoughi' 'Nikolay Yakovets' 'George Fletcher'] The ISO standard Property Graph model has become increasingly popular for representing complex, interconnected data. However, it lacks native support for querying metadata and reification, which limits its abilities to deal with the…
Darko Hric, Tiago P. Peixoto, Santo Fortunato
The empirical validation of community detection methods is often based on available annotations on the nodes that serve as putative indicators of the large-scale network structure. Most often, the suitability of the annotations as topological descriptors itself is not assessed, and without this it is not possible to…
Vladimir Batagelj, Tomaž Pisanski, Iztok Savnik, Ana Slavec + 1 more
- Python: NetworkX, igraph, Snap.py, graph-tool, NetworKit, PyGraphistry, Nets, cdlib, node2vec, DGL, PyG, Tulip, PyVis, - R: igraph, statnet, sna, qgraph, RSiena, tnet, multiplex, NetSim, influenceR, tidygraph, intergraph, netUtils, ggraph, networkD3, visNetwork, DiagrammeR, graphlayouts, ndtv, - Julia: LightGraphs…
Lena Mangold, Camille Roth
Network analysis is often enriched by including an examination of node metadata. In the context of understanding the mesoscale of networks it is often assumed that node groups based on metadata and node groups based on connectivity patterns are intrinsically linked. This assumption is increasingly being challenged…
Spyros Blanas, Suren Byna
Advances in technology and computing hardware are enabling scientists from all areas of science to produce massive amounts of data using large-scale simulations or observational facilities. In this era of data deluge, effective coordination between the data production and the analysis phases hinges on the availability…
Pável Vázquez, Kayoko Shoji, Steffen Nøvik, Stefan Krauß + 1 more
'Simon Rayner'] GADDS: Global Accessible Distribution Data Sharing Machine: physical hardware that can execute commands. Node: machine in a network. Cluster: group of machines. Organization: group of nodes sharing a domain name. Domain: network address. Channel: permissioned network where organizations communicate.…
Salman Niazi, Mahmoud Ismail, Seif Haridi, Jim Dowling
Recent improvements in both the performance and scalability of shared-nothing, transactional, in-memory NewSQL databases have reopened the research question of whether distributed metadata for hierarchical file systems can be managed using commodity databases. In this paper, we introduce HopsFS, a next generation…
Donald Pinckney, Federico Cassano, Arjun Guha, Jonathan Bell
Software developers typically rely upon a large network of dependencies to build their applications. For instance, the NPM package repository contains over 3 million packages and serves tens of billions of downloads weekly. Understanding the structure and nature of packages, dependencies, and published code requires…
Robert Primmer, Scott Nyman, Wayzen Lin
In and of itself, data storage has apparent business utility. But when we can convert data to information, the utility of stored data increases dramatically. It is the layering of relation atop the data mass that is the engine for such conversion. Frank relation amongst discrete objects sporadically ingested is rare…
Simone Bocca, Amarsanaa Ganbold, Tsolmon Zundui
—Data reuse is fundamental for reducing the data integration effort required to build data supporting new applications, especially in data scarcity contexts. However, data reuse requires to deal with data heterogeneity, which is always present in data coming from different sources. Such heterogeneity appears at…
Florian Rupp, Benjamin Schnabel, Kai Eckert
The Resource Description Framework is well-established as a lingua franca for data modeling and is designed to integrate heterogeneous data at instance and schema level using statements. While RDF is conceptually simple, data models nevertheless get complex, when complex data needs to be represented. Additional levels…