Search · four archives
Search · four archives
21 papers · ranked by Valyu relevance
Adeyinka Akanbi, Muthoni Masinde
In recent years, the application and wide adoption of Internet of Things (IoT)-based technologies have increased the proliferation of monitoring systems, which has consequently exponentially increased the amounts of heterogeneous data generated. Processing and analysing the massive amount of data produced is cumbersome…
Teng Li, Zhiyuan Xu, Jian Tang, Yanzhi Wang
In this paper, we focus on general-purpose Distributed Stream Data Processing Systems (DSDPSs), which deal with processing of unbounded streams of continuous data at scale distributedly in real or near-real time. A fundamental prob lem in a DSDPS is the scheduling problem (i.e., assigning workload to workers/machines)…
Muhammad Hanif, Choonhwa Lee, Sumi Helal, Rashid Mehmood
Cloud computing has evolved the big data technologies to a consolidated paradigm with SPaaS (Streaming processing-as-a-service). With a number of enterprises offering cloud-based solutions to end-users and other small enterprises, there has been a boom in the volume of data, creating interest of both industry and…
Marcos Dias de Assunção, Alexandre da Silva Veith, Rajkumar Buyya
Under several emerging application scenarios, such as in smart cities, operational monitoring of large infrastructure, wearable assistance, and Internet of Things, continuous data streams must be processed under very short delays. Several solutions, including multiple software engines, have been developed for…
Tarek Stolz, István Koren, Liam Tirpitz, Sandra Geisler
With the increasing prevalence of IoT environments, the demand for processing massive distributed data streams has become a critical challenge. Data Stream Processing on the Edge (DSPoE) systems have emerged as a solution to address this challenge, but they often struggle to cope with the heterogeneity of hardware and…
Naoual El aboudi, Laila Benhlima
The growing amount of data in healthcare industry has made inevitable the adoption of big data techniques in order to improve the quality of healthcare delivery. Despite the integration of big data processing approaches and platforms in existing data management architectures for healthcare systems, these architectures…
Xikui Wang, Michael J. Carey, Vassilis J. Tsotras
Today, data is being actively generated by a variety of devices, services, and applications. Such data is important not only for the information that it contains, but also for its relationships to other data and to interested users. Most existing Big Data systems focus on passively answering queries from users, rather…
Francesco Versaci, Luca Pireddu, Gianluigi Zanetti
Modern sequencing machines produce order of a terabyte of data per day, which need subsequently to go through a complex processing pipeline. The standard workflow begins with a few independent, shared-memory tools, which communicate by means of intermediate files. Given the constant increase of the amount of data…
Nirav Bhatt, Amit Thakkar, Othman Soufan
Stream data is the data that is generated continuously from the different data sources and ideally defined as the data that has no discrete beginning or end. Processing the stream data is a part of big data analytics that aims at querying the continuously arriving data and extracting meaningful information from the…
Zhijian Qu, Hanxin Liu, Hanlin Wang, Xinqiang Chen + 3 more
'Zixiao Wang' 'Sakdirat Kaewunruen'] The purpose of the study is to solve problems, i.e., increasingly significant processing delay of massive monitoring data and imbalanced tasks in the scheduling and monitoring center for a railway network. To tackle these problems, a method by using a smooth weighted round-robin…
Ben Blamey, Salman Toor, Martin Dahlö, Håkan Wieslander + 6 more
This paper introduces the HASTE Toolkit, a cloud-native software toolkit capable of partitioning data streams in order to prioritize usage of limited resources. This in turn enables more efficient data-intensive experiments. We propose a model that introduces automated, autonomous decision making in data pipelines…
Shantenu Jha, Daniel S. Katz, André Luckow, Omer Rana + 2 more
'Yogesh Simmhan' 'Neil Chue Hong'] A common feature across many science and engineering applications is the amount and diversity of data and computation that must be integrated to yield insights. Data sets are growing larger and becoming distributed; and their location, availability and properties are often…
Ovidiu-Cristian Marcu, Pascal Bouvry
—Real-time Big Data architectures evolved into specialized layers for handling data streams' ingestion, storage, and processing over the past decade. Layered streaming architectures integrate pull-based read and push-based write RPC mechanisms implemented by stream ingestion/storage systems. In addition, stream…
Cristian Ramon-Cortes, Francesc Lordan, Jorge Ejarque, Rosa M. Badía
In the past years, e-Science applications have evolved from large-scale simulations executed in a single cluster to more complex workflows where these simulations are combined with High-Performance Data Analytics (HPDA). To implement these workflows, developers are currently using different patterns; mainly task-based…
Authors not listed
Machine learning models are transforming data-driven research across scientific disciplines, yet their deployment as accessible and reliable web services remains a significant challenge. We introduce the NERDD framework, a scalable, maintainable, and secure microservices platform designed to support the sustainable…
Tanveer Ahmad, Chengxin Ma, Zaid Al-Ars, H. Peter Hofstee
Current cluster scaled genomics data processing solutions rely on big data frameworks like Apache Spark, Hadoop and HDFS for data scheduling, processing and storage. These frameworks come with additional computation and memory overheads by default. It has been observed that scaling genomics dataset processing beyond 32…
Sylvain Hallé, Raphaël Khoury, Sébastien Gaboury
Current runtime verification tools seldom make use of multi-threading to speed up the evaluation of a property on a large event trace. In this paper, we present an extension to the BeepBeep 3 event stream engine that allows the use of multiple threads during the evaluation of a query. Various parallelization strategies…
J.J. Salvo, M. Lakshman, A.M. Holubecki, Z.M. Saygin + 2 more
Reading bridges sensation and cognition. To derive meaning from written words, visual input is first processed in unimodal (i.e., sensory-specific) visual streams and then engages a distributed language network (LANG) that includes classic perisylvian language areas and supports transmodal (i.e., sensory-nonspecific)…
Binke Yuan, Hui Xie, Zhihao Wang, Yangwen Xu + 11 more
Modern linguistic theories and network science propose that the language and speech processing is organized into hierarchical, segregated large-scale subnetworks, with a core of dorsal (phonological) stream and ventral (semantic) stream. The two streams are asymmetrically recruited in receptive and expressive language…
Benjamin Coleman, Benito Geordie, Li Chou, R. A. Leo Elworth + 2 more
The rise of whole-genome shotgun sequencing (WGS) has enabled numerous breakthroughs in large-scale comparative genomics research. However, the size of genomic datasets has grown exponentially over the last few years, leading to new challenges for traditional streaming algorithms. Modern petabyte-sized genomic datasets…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…