Search · four archives
Search · four archives
15 papers · ranked by Valyu relevance
Spenger, Jonas, Krafeld, Kolya + 6 more
Scaling global aggregations is a challenge for exactly-once stream processing systems. Current systems implement these either by computing the aggregation in a single task instance, or by static aggregation trees, which limits scalability and may become a bottleneck. Moreover, the end-to-end latency is determined by…
Hung T. Pham, Viet Vo, Tien Tuan Anh Dinh, Duc H. Tran + 1 more
—Stream processing systems are important in modern applications in which data arrive continuously and need to be processed in real time. Because of their resource and scalability requirements, many of these systems run on the cloud, which is considered untrusted. Existing works on securing databases on the cloud focus…
Ander Cejudo, Yone Tellechea, Amaia Calvo, Aitor Almeida + 3 more
Background The increasing use of real-time health data from wearable devices and self-reported questionnaires offers significant opportunities for preventive care in aging populations. However, current health data platforms often lack built-in mechanisms for data and model traceability, version control, and coordinated…
G.P. Saggese, Paul Smith
| 1. | Introduction | 1 | |----|--------------------------------------------|----| | 2. | DataFlow at a Glance | 6 | | 3. | Challenges in time series machine learning | 7 | | 4. | Semantics | 11 | | 5. | DAGs | 29 | | 6. | Execution Engine | 38 | | 7. | Comparison to Related Work | 47 | | | References | 49 |
Cao, Jiaping, Ting Sun, Man Lung Yiu + 2 more
—Spatial data analytics systems are widely studied in both the academia and industry. However, existing systems are limited when handling a large number of moving objects and realtime spatial queries. In this work, we architect a scalable and efficient system CheetahGIS to process streaming spatial queries over massive…
Monica Marconi Sciarroni, Emanuele Storti
Industrial IoT ecosystems bring together sensors, machines and smart devices operating collaboratively across industrial environments. These systems generate large volumes of heterogeneous, high-velocity data streams that require interoperable, secure and contextually aware management. Most of the current stream…
Kyriakos Psarakis, George Christodoulou, George Siachamis, Marios Fragkoulis + 1 more
Developing stateful cloud applications, such as low-latency workflows and microservices with strict consistency requirements, remains arduous for programmers. The Stateful Functions-as-a-Service (SFaaS) paradigm aims to serve these use cases. However, existing approaches provide weak transactional guarantees or perform…
Peter G. Hawkins, Eli M. Swanson, Megan Feichtel
The size of individual single cell samples continues to grow with advancing technologies, as do the number of samples included in individual experiments and across organizations. This presents challenges for processing this data at scale, both in terms of computational throughput and the required size of the machines…
Nico Migenda, Ralf Möller, Wolfram Schenck, Muhammad Ahsan
We present H-NGPCA, a hierarchical clustering algorithm for data streams that integrates an adaptive unit number growth and local dimensionality control. Unlike existing algorithm, H-NGPCA combines the characteristics of centroid-based, model-based and hierarchical clustering. H-NGPCA builds a hierarchical structure of…
Mahmudur Rahman Hera, David Koslicki, Conrado Martínez
With the surge in sequencing data generated from an ever-expanding range of biological studies, designing scalable computational techniques has become essential. One effective strategy to enable large-scale computation is to split long DNA or protein sequences into k-mers, and summarize large k-mer sets into compact…
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Thang V Pham, Chau TM Tran, Alex A Henneman, Long HC Pham + 5 more
Current methods for protein level quantification in mass spectrometry-based proteomics do not scale with the increasing number of samples because of limited system memory and algorithmic complexities. Here we propose a new data structure that supports parsing of input as data stream, improve state of the art…
Phanindra Reddy Madduru, Bijo Thomas
This paper proposes a preprocessing framework for optimizing large-scale graph database ingestion through intelligent edge filtering based on value ranking. We combine adapted PageRank algorithms with business-specific metrics and edge type importance to evaluate and rank edges, enabling selective retention of…
Kevin Garner, Polykarpos Thomadakis, Nikos Chrisochoides
This paper presents a distributed memory method for anisotropic mesh adaptation that is designed to avoid the use of collective communication and global synchronization techniques. In the presented method, meshing functionality is separated from performance aspects by utilizing a separate entity for each - a multicore…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…