17 papers · ranked by Valyu relevance
Duling Xu, Tong Li, Zegang Sun, Zheng Chen + 4 more
The deployment of databases across geographically distributed regions has become increasingly critical for ensuring data reliability and scalability. Recent studies indicate that distributed databases exhibit significantly higher latency than single-node databases, primarily due to consensus protocols maintaining data…
Şenol Zaman
—Modern cloud databases present scaling as a binary decision: scale-out by adding nodes or scale-up by increasing pernode resources. This one-dimensional view is limiting because database performance, cost, and coordination overhead emerge from the joint interaction of horizontal elasticity and per-node CPU, memory…
Oto Mraz, Kyriakos Psarakis, George Christodoulou, Paris Carbone + 1 more
Geo-distributed OLTP databases are widely deployed across cloud regions, yet current evaluation practices do not cover the challenges of this aspect. Existing benchmarks assume stable network conditions; they lack explicit settings for data and client locality, and they largely ignore data transfer costs across…
Raaghav Ravishankar, Sandeep S. Kulkarni, Nitin H. Vaidya
We focus on the problem of checkpointing (or snapshotting) in fully replicated weakly consistent distributed databases, which we refer to as Distributed Transaction Consistent Snapshots (DTCS). A typical example of such a system is a replicated main-memory database that provides strong eventual consistency. This…
Nouf Aljuaid, Alexei Lisitsa, Sven Schewe
We have developed a framework for efficient privacy preserving multi-party querying (PPMQ) over federated graph databases, leveraging Secure Multi-Party Computation (SMPC) protocols to enhance data security. The system offers two distinct security protocols: a client-based protocol and a server-based protocol. In the…
Authors not listed
Next Generation Risk Assessment (NGRA) promotes animal-free, exposure-informed, and hypothesis-driven approaches to chemical safety assessment. In silico tools, such as quantitative structure-activity relationship (QSAR) models, are valuable new approach methodologies (NAMs) for use in NGRA. However, the practical…
Heidi J. Imker
For decades, life science researchers have had cost-free, unrestricted access to data through online databases. However, the sustainability of even well-established resources was already tenuous, and abrupt changes in science funding in the United States seems poised to exacerbate these challenges. This study employed…
Roy Shadmon, Mark Davidson, Eric Aquaronne, Massimiliano Pinto + 2 more
Industrial and autonomous systems increasingly depend on AI, automation, and real-time coordination to act on operational data as it is generated. Yet conventional architectures often require that data to pass through centralized platforms before decisions can be made. Cloud systems remain valuable for training…
Chen Chen, Yuanyuan Liu, Lei Wang, Jingyi Sai + 8 more
With the rapid accumulation of diverse omics datasets, achieving efficient management and integrative analysis of plant multi-omics data remains a major challenge. Conventional solutions rely on constructing web-based databases, which often demand substantial programming expertise and long-term financial support. To…
Fabrizio Carinci, Stephen Fava, Iztok Štotl, Nicholas Nicholson
Background The growing burden of non-communicable diseases (NCDs) in Europe has intensified the need for timely, comparable, and policy-relevant health indicators derived from increasingly heterogeneous health data ecosystems. The European Health Data Space (EHDS) represents a major policy initiative to facilitate the…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Xiaowen Suo, Fuzhong Xue, Yanyan Zhao
Genome-wide association studies (GWAS) increasingly rely on large-scale data integration to achieve the statistical power necessary to detect variants with weak effects. However, genomic data are typically siloed across institutions, and privacy constraints often preclude centralized analysis. While federated learning…
Chun Yang, Yining Ma
This paper proposes a secure and efficient data aggregation approach for heterogeneous enterprise data, leveraging federated meta-learning (FML) and data consolidation. The proposed approach addresses critical challenges in enterprise data, including data privacy, heterogeneous data distributions, and communication…
Peter G. Hawkins, Eli M. Swanson, Megan Feichtel
The size of individual single cell samples continues to grow with advancing technologies, as do the number of samples included in individual experiments and across organizations. This presents challenges for processing this data at scale, both in terms of computational throughput and the required size of the machines…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Paul L. Soto
Data collection and analysis are central to scientific research, including in applied and basic behavior analysis. A substantial amount of attention has been given to how to rigorously collect and analyze data. Less attention has been paid to storing and maintaining research data, which becomes a critical step in the…
Miklós Bán
The increasing reliance of biodiversity research on large-scale databases has brought significant progress in data accessibility but also new challenges in data comparability, reliability, and interpretability. While global platforms, such as GBIF and iNaturalist standardize and disseminate vast quantities of…