21 papers · ranked by Valyu relevance
Danila Valko, Jorge Marx Gómez
The Lightning Network (LN) is the most widely adopted second-layer solution for Bitcoin, enabling fast, low-cost transactions through a decentralized payment channel network. Despite its growing importance and the increasing interest from researchers across disciplines, progress in LN research is often impeded by…
Junru Li, Qing Wang, Zhe Yang, Shuo Liu + 2 more
—Distributed storage systems typically maintain strong consistency between data nodes and metadata nodes by adopting ordered writes: 1) first installing data; 2) then updating metadata to make data visible. We propose Switch∆ to accelerate ordered writes by moving metadata updates out of the critical path. It buffers…
Andrews Frimpong Adu, Elliot Sarpong Menkah, Peter Amoako-Yirenkyi, Samson Pandam Salifu
De novo genome assembly using de Bruijn graphs (DBGs) typically relies on fixed-length $k$-mers as the nodes of the graph. While this approach is effective, it presents a fundamental trade-off: smaller $k$ values tend to collapse repeats, whereas larger $k$ values can result in fragmentation, particularly in…
Yahya Sa'd, Renzo Angles, Vojtech Merunka, Roberto Garcia + 2 more
Property-graph schemas often contain descriptive properties that recur across heterogeneous nodes and edges, yet schema designers lack a clear method for deciding whether such properties should remain embedded or be treated as reusable metadata structures. This paper addresses this design-stage problem within a…
F. Kirchner, D.C.F. Wieland, S. Irvine, S. Schimek + 11 more
Scientific communities have recognized the importance of well-documented metadata generated during research. However, ensuring that metadata is findable, accessible, interoperable, and reusable (FAIR) remains a significant challenge. To address this, scientific communities are working towards making metadata available…
Richard Arthur, Virginia DiDomizio, Louis Hoebel
In some complex domains, certain problem-specific decompositions can provide advantages over monolithic designs by enabling comprehension and specification of the design. In this paper we present an intuitive and tractable approach to reasoning over large and complex data sets. Our approach is based on Active Data…
Jasmin Walter, Carsten Kuenne, Noah Knoppik, Philipp Goymann + 1 more
Scientific research relies on transparent dissemination of data and its associated interpretations. This task encompasses accessibility of raw data, its metadata, details concerning experimental design, along with parameters and tools employed for data interpretation. Production and handling of these data represents an…
Polina Shpilker, Benjamin J. Stubbs, Michael W. Sayers, Yumin Lee + 5 more
Scientific research metadata is vital to ensure the validity, reusability, and cost-effectiveness of research efforts. The MEDFORD metadata language was previously introduced to simplify the process of writing and maintaining metadata for non-programmers. However, barriers to entry and usability remain, including…
Qing Wu, Ailing Zhang, Zhibin Ning, Daniel Figeys
Hierarchically structured data are common in biology but are difficult to visualize in a way that both preserves structure and supports systematic comparison of many samples and experimental groups. Conventional approaches either discard the hierarchy or focus on single annotated trees, making it difficult to…
Dongyang Fan, Diba Hashemi, Sai Praneeth Karimireddy, Martin Jaggi
Incorporating metadata in Large Language Models (LLMs) pretraining has recently emerged as a promising approach to accelerate training. However prior work highlighted only one useful signal—URLs, leaving open the question of whether other forms of metadata could yield greater benefits. In this study, we investigate a…
Michail Lazaratos, Neele Haacke, Jasmin Gaugel, Miriam Ulz + 4 more
metaKEGG is a comprehensive software package designed to streamline the visualization and integration of pathway enrichment results from multi-omics data, providing accessible and detailed insights into the molecular mechanisms driving health and disease. Unlike standard pipeline approaches, metaKEGG incorporates novel…
Fabrizio Musacchio, Henrike Antony, Arush Baijal, Sophie Crux + 7 more
Modern fluorescence and multiphoton microscopy workflows operate within a heterogeneous ecosystem of file formats, partially overlapping metadata standards, and reader-specific conventions. In practice, this frequently leads to silent axis misinterpretations, loss or corruption of physical voxel size information, and…
Harald Witte, Deepak Raveendran Unni, Philip Krauss, Vasundra Touré + 3 more
Background The Swiss Personalized Health Network (SPHN) facilitates the interoperability and secure sharing of health-related data for research in Switzerland, in line with the findable, accessible, interoperable, and reusable (FAIR) principles. Since medical datasets can be highly sensitive, access is often governed…
Liang Zhang, Xin Lai
The exponential growth of data in biomedicine has created an urgent need for intuitive visualization tools. These tools must be able to effectively represent complex biological networks and remain accessible to domain experts without extensive computational training. Current network visualization approaches often…
Huanfei Wang, Shixue Sun, Ewy A. Mathé, Qian Zhu
Rare diseases (RD) impact over 30 million individuals in the United States, yet fewer than 5% of the identified conditions have FDA-approved treatments. Progress in RD research is hindered by small patient cohorts, biological heterogeneity, and the fragmented, inconsistently annotated publicly available omics data…
Chunyu Ma, Shaopeng Liu, Stephanie Won, David Koslicki + 1 more
Our MetagenomicKG integrates data from seven sources: GTDB taxonomy (), NCBI taxonomy (), KEGG (), RTX-KG2 (), BV-BRC (), MicroPhenoDB (), and NCBI AMRFinderPlus Prediction () (see [btag421-F1]). MetagenomicKG is a directed multigraph comprising 1.25 million nodes and 56 million edges. Nodes are categorized into 14…
Dorothea Strecker
This paper investigates how eight disciplinary research data repositories from the geosciences and social sciences navigate metadata conflicts - conflicts in implementations of the same standard and inter-standard conflicts - and how these conflicts affect the completeness of DataCite metadata. It combines results from…
Authors not listed
Recent advances in generative artificial intelligence have enabled in silico molecular design to become a powerful approach for exploring chemical space toward specific design goals across various domains. However, in actual design workflows, determining the appropriate generation conditions, including generative…
Authors not listed
Agentic artificial intelligence (AI) is poised to redefine how science is conducted, automating not just data analysis but the entire research lifecycle, from hypothesis generation to validation. Yet most current AI agents remain domain-bound, tailored to specific applications such as materials synthesis or quantum…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…