18 papers · ranked by Valyu relevance
Larissa Mori, Kaleigh O’Hara, Toyya A. Pujol, Mario Ventresca + 1 more
With the goal of understanding if the information contained in node metadata can help in the task of link weight prediction, we investigate herein whether incorporating it as a similarity feature (referred to as metadata similarity) between end nodes of a link improves the prediction accuracy of common supervised…
James P. Bagrow, Yong‐Yeol Ahn
The deluge of network datasets demands a standard way to effectively and succinctly summarize network datasets. Building on similar efforts to standardize the documentation of models and datasets in machine learning, here we propose network cards, short summaries of network datasets that can capture not only the basic…
Paul Billing Ross, Jina Song, Philip S. Tsao, Cuiping Pan
Biomedical studies have become larger in size and yielded large quantities of data, yet efficient data processing remains a challenge. Here we present Trellis, a cloud-based data and task management framework that completely automates the process from data ingestion to result presentation, while tracking data lineage…
Aleix Bassolas, Anton Holmgren, Antoine Marot, Martin Rosvall + 1 more
'Vincenzo Nicosia'] Integrating structural information and metadata, such as gender, social status, or interests, enriches networks and enables a better understanding of the large-scale structure of complex systems. However, existing approaches to augment networks with metadata for community detection only consider…
Sepehr Sadoughi, Nikolay Yakovets, George Fletcher
and Reification Authors: ['Sepehr Sadoughi' 'Nikolay Yakovets' 'George Fletcher'] The ISO standard Property Graph model has become increasingly popular for representing complex, interconnected data. However, it lacks native support for querying metadata and reification, which limits its abilities to deal with the…
Yahya Sa'd, Renzo Angles, Vojtech Merunka, Roberto Garcia + 2 more
Property-graph schemas often contain descriptive properties that recur across heterogeneous nodes and edges, yet schema designers lack a clear method for deciding whether such properties should remain embedded or be treated as reusable metadata structures. This paper addresses this design-stage problem within a…
Taha Mohseni Ahooyi, Benjamin Stear, J. Alan Simmons, Vincent T. Metzger + 40 more
The Data Distillery Knowledge Graph (DDKG) is a framework for semantic integration and querying of biomedical data across domains. Built for the NIH Common Fund Data Ecosystem, it supports translational research by linking clinical and experimental datasets in a unified graph model. Clinical standards such as ICD-10…
Authors not listed
Effective visualization of complex synthesis routes is critical for computeraided synthesis planning (CASP), yet current solutions are limited in scope, integration flexibility, and chemical intuition. We introduce RouteWise, a versatile, containerized web application designed to address these unmet needs. Its modular…
Lena Mangold, Camille Roth
Network analysis is often enriched by including an examination of node metadata. In the context of understanding the mesoscale of networks it is often assumed that node groups based on metadata and node groups based on connectivity patterns are intrinsically linked. This assumption is increasingly being challenged…
Chunyu Ma, Shaopeng Liu, David Koslicki
The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto…
Donald Pinckney, Federico Cassano, Arjun Guha, Jonathan Bell
Software developers typically rely upon a large network of dependencies to build their applications. For instance, the NPM package repository contains over 3 million packages and serves tens of billions of downloads weekly. Understanding the structure and nature of packages, dependencies, and published code requires…
Simone Bocca, Amarsanaa Ganbold, Tsolmon Zundui
—Data reuse is fundamental for reducing the data integration effort required to build data supporting new applications, especially in data scarcity contexts. However, data reuse requires to deal with data heterogeneity, which is always present in data coming from different sources. Such heterogeneity appears at…
Daniel S. Himmelstein, Michael Zietz, Vincent Rubinetti, Kyle Kloster + 8 more
Hetnets, short for “heterogeneous networks”, contain multiple node and relationship types and offer a way to encode biomedical knowledge. One such example, Hetionet connects 11 types of nodes - including genes, diseases, drugs, pathways, and anatomical structures - with over 2 million edges of 24 types. Previous work…
Nolan K. Newman, Matthew Macovsky, Richard R. Rodrigues, Amanda M. Bruce + 8 more
Technological advances have generated tremendous amounts of high-throughput omics data. Integrating data from multiple cohorts and/or several omics types from new and previously published studies can offer a holistic view of a biological system and decipher its critical players and key mechanisms. In this protocol we…
Daniel S. Himmelstein, Michael Zietz, Vincent Rubinetti, Kyle Kloster + 8 more
Hetnets, short for “heterogeneous networks”, contain multiple node and relationship types and offer a way to encode biomedical knowledge. One such example, Hetionet connects 11 types of nodes — including genes, diseases, drugs, pathways, and anatomical structures — with over 2 million edges of 24 types. Previous work…
Dongjun Na, Jinbum Kim, Juseong Jeon, Sejin Park + 3 more
'Victor C.M. Leung' 'Sara Rouhani'] Blockchain technology can address data falsification, single point of failure (SPOF), and DDoS attacks on centralized services. By utilizing IoT devices as blockchain nodes, it is possible to solve the problem that it is difficult to ensure the integrity of data generated by using…
Zhenyu Wei, Chengkui Zhao, Min Zhang, Jiayu Xu + 5 more
'Xiaohui Xin' 'Lei Yu' 'Weixing Feng'] Title: Abstract Chimeric antigen receptor T-cell (CAR-T) immunotherapy, a novel approach for treating blood cancer, is associated with the production of cytokine release syndrome (CRS), which poses significant safety concerns for patients. Currently, there is limited knowledge…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…