19 papers · ranked by Valyu relevance
Matteo Cinelli, Giovanna Ferraro, Antonio Iovanella
Networks are real systems modelled through mathematical objects made up of nodes and links arranged into peculiar and deliberate (or partially deliberate) topologies. Studying these real-world topologies allows for several properties of interest to be revealed. In real networks, nodes are also identified by a certain…
Matteo Cinelli, Leto Peel, Antonio Iovanella, Jean‐Charles Delvenne
We consider the network constraints on the bounds of the assortativity coefficient, which aims to quantify the tendency of nodes with the same attribute values to be connected. The assortativity coefficient can be considered as the Pearson's correlation coefficient of node metadata values across network edges and lies…
Aleix Bassolas, Anton Holmgren, Antoine Marot, Martin Rosvall + 1 more
'Vincenzo Nicosia'] Integrating structural information and metadata, such as gender, social status, or interests, enriches networks and enables a better understanding of the large-scale structure of complex systems. However, existing approaches to augment networks with metadata for community detection only consider…
Gen-Tao Chiang, Peter Clapham, Guoying Qi, Kevin Sale + 1 more
Background Increasingly large amounts of DNA sequencing data are being generated within the Wellcome Trust Sanger Institute (WTSI). The traditional file system struggles to handle these increasing amounts of sequence data. A good data management system therefore needs to be implemented and integrated into the current…
Jesper Nielsen, Thomas Mailund
Background High-throughput genotyping technology has enabled cost effective typing of thousands of individuals in hundred of thousands of markers for use in genome wide studies. This vast improvement in data acquisition technology makes it an informatics challenge to efficiently store and manipulate the data. While…
Daniela Ballari, Monica Wachowicz, Miguel Angel Manso Callejo
Wireless Sensor Networks (WSNs) produce changes of status that are frequent, dynamic and unpredictable, and cannot be represented using a linear cause-effect approach. Consequently, a new approach is needed to handle these changes in order to support dynamic interoperability. Our approach is to introduce the notion of…
Lucas Czech, Jaime Huerta-Cepas, Alexandros Stamatakis
Phylogenetic trees are routinely visualized to present and interpret the evolutionary relationships of species. Virtually all empirical evolutionary data studies contain a visualization of the inferred tree with branch support values. Ambiguous semantics in tree file formats can lead to erroneous tree visualizations…
Basavaraja Bheemalingappa Sagar, Georg von Hippel, Giannis Koutsou, Hideo Matsufuru + 3 more
- 𝑑Computing Research Center, High Energy Accelerator Research Organization (KEK), and Accelerator Science Program, Graduate Institute for Advanced Studies, Graduate University for Advanced Studies (SOKENDAI), Oho 1-1, Tsukuba 305-0801, Japan - 𝑒Bernoulli Institute for Mathematics, Computer Science and Artificial…
Peng Sun, Yonggang Wen, Duong Nguyen Binh Ta, Haiyong Xie
—In large-scale distributed file systems, efficient metadata operations are critical since most file operations have to interact with metadata servers first. In existing distributed hash table (DHT) based metadata management systems, the lookup service could be a performance bottleneck due to its significant CPU…
Noam Teyssier, Alexander Dobin
Modern genomics produces billions of sequencing records per run, which are typically stored as gzip-compressed FASTQ files. While this format is widely used, it is not optimal for high-throughput processing due to its reliance on single-threaded decompression and sequential parsing of irregularly sized records. This…
Authors not listed
Effective visualization of complex synthesis routes is critical for computeraided synthesis planning (CASP), yet current solutions are limited in scope, integration flexibility, and chemical intuition. We introduce RouteWise, a versatile, containerized web application designed to address these unmet needs. Its modular…
F. Kirchner, D.C.F. Wieland, S. Irvine, S. Schimek + 11 more
Scientific communities have recognized the importance of well-documented metadata generated during research. However, ensuring that metadata is findable, accessible, interoperable, and reusable (FAIR) remains a significant challenge. To address this, scientific communities are working towards making metadata available…
Lena Mangold, Camille Roth
Network analysis is often enriched by including an examination of node metadata. In the context of understanding the mesoscale of networks it is often assumed that node groups based on metadata and node groups based on connectivity patterns are intrinsically linked. This assumption is increasingly being challenged…
Mahnoor Zulfiqar, Michael R. Crusoe, Birgitta König-Ries, Christoph Steinbeck + 2 more
Scientific workflows facilitate the automation of data analysis tasks by integrating various software and tools executed in a particular order. To enable transparency and reusability in workflows, it is essential to implement the FAIR principles. Here, we describe our experiences implementing the FAIR principles for…
Hadi Poormohammadi, Mohsen Sardari Zarchi
Phylogenetic networks construction is one the most important challenge in phylogenetics. These networks can present complex non-treelike events such as gene flow, horizontal gene transfers, recombination or hybridizations. Among phylogenetic networks, rooted structures are commonly used to represent the evolutionary…
Thurston H. Y. Dang, José Cambronero, Martin Rinard
We present BIEBER (Byte-IdEntical Binary parsER), the first system to model and regenerate a full working parser from instrumented program executions. To achieve this, BIEBER exploits the regularity (e.g., header fields and array-like data structures) that is commonly found in file formats. Key generalization steps…
Rongjie Wang, Junyi Li, Yang Bai, Tianyi Zang + 2 more
Dramatic increases in data produced by next-generation sequencing (NGS) technologies demand data compression tools for saving storage space. However, effective and efficient data compression for genome sequencing data has remained an unresolved challenge in NGS data studies. In this paper, we propose a novel…
Lijia Jia, Yue Shi, Jing Yang, Shangzhe Li + 14 more
The explosive growth of digital data is overwhelming conventional storage media, creating an urgent need for more efficient solutions. DNA offers immense potential for digital data storage, yet most systems remain static and archival. Here, we present a modular DNA storage architecture based on dynamic DNA bytes…
Heng Li, Jiazhen Rong
We present bedtk, a new toolkit for manipulating genomic intervals in the BED format. It supports sorting, merging, intersection, subtraction and the calculation of the breadth of coverage. Bedtk employs implicit interval tree, a new data structure for fast interval overlap queries. It is several to tens of times…