20 papers · ranked by Valyu relevance
Klemen Kenda, Blaž Kažič, Erik Novak, Dunja Mladenić
To achieve the full analytical potential of the streaming data from the internet of things, the interconnection of various data sources is needed. By definition, those sources are heterogeneous and their integration is not a trivial task. A common approach to exploit streaming sensor data potential is to use machine…
Genoveva Vargas‐Solar, Javier A. Espinosa-Oviedo
This paper introduces H-STREAM, a big stream/data processing pipelines evaluation engine that proposes stream processing operators as micro-services to support the analysis and visualisation of Big Data streams stemming from IoT (Internet of Things) environments. H-STREAM micro-services combine stream processing and…
Indrė Žliobaitė, Jesse Read
Machine learning from data streams is an active and growing research area. Research on learning from streaming data typically makes strict assumptions linked to computational resource constraints, including requirements for stream mining algorithms to inspect each instance not more than once and be ready to give a…
Alaettin Zubaroğlu, Volkan Atalay
Number of connected devices is steadily increasing and these devices continuously generate data streams. Real-time processing of data streams is arousing interest despite many challenges. Clustering is one of the most suitable methods for real-time data stream processing, because it can be applied with less prior…
Haruna Isah, Farhana Zulkernine
—An essential part of building a data-driven organization is the ability to handle and process continuous streams of data to discover actionable insights. The explosive growth of interconnected devices and the social Web has led to a large volume of data being generated on a continuous basis. Streaming data sources…
Adeyinka Akanbi, Muthoni Masinde
In recent years, the application and wide adoption of Internet of Things (IoT)-based technologies have increased the proliferation of monitoring systems, which has consequently exponentially increased the amounts of heterogeneous data generated. Processing and analysing the massive amount of data produced is cumbersome…
GonÇalo Lopes, Niccolò Bonacchi, João Frazão, Joana P. Neto + 11 more
The design of modern scientific experiments requires the control and monitoring of many parallel data streams. However, the serial execution of programming instructions in a computer makes it a challenge to develop software that can deal with the asynchronous, parallel nature of scientific data. Here we present Bonsai…
Ayush Singhal, Rakesh Pant, Pradeep K. Sinha
The demand for stream processing is increasing at an unprecedented rate. Big data is no longer limited to processing of big volumes of data. In most real-world scenarios, the need for processing stream data as it comes can only meet the business needs. It is required for trading, fraud detection, system monitoring…
Xikui Wang, Michael J. Carey, Vassilis J. Tsotras
Today, data is being actively generated by a variety of devices, services, and applications. Such data is important not only for the information that it contains, but also for its relationships to other data and to interested users. Most existing Big Data systems focus on passively answering queries from users, rather…
Naoual El aboudi, Laila Benhlima
The growing amount of data in healthcare industry has made inevitable the adoption of big data techniques in order to improve the quality of healthcare delivery. Despite the integration of big data processing approaches and platforms in existing data management architectures for healthcare systems, these architectures…
Adeyinka K. Akanbi
Distributed networks and real-time systems are becoming the most important components for the new computer age – the Internet of Things (IoT), with huge data streams or data sets generated from sensors and data generated from existing legacy systems. The data generated offers the ability to measure, infer and…
Marios Fragkoulis, Paris Carbone, Vasiliki Kalavri, Asterios Katsifodimos
'Asterios Katsifodimos'] Abstract Stream processing has been an active research field for more than 20 years, but it is now witnessing its prime time due to recent successful efforts by the research community and numerous worldwide open-source communities. This survey provides a comprehensive overview of fundamental…
Ben Blamey, Salman Toor, Martin Dahlö, Håkan Wieslander + 6 more
This paper introduces the HASTE Toolkit, a cloud-native software toolkit capable of partitioning data streams in order to prioritize usage of limited resources. This in turn enables more efficient data-intensive experiments. We propose a model that introduces automated, autonomous decision making in data pipelines…
Benjamin Coleman, Benito Geordie, Li Chou, R. A. Leo Elworth + 2 more
The rise of whole-genome shotgun sequencing (WGS) has enabled numerous breakthroughs in large-scale comparative genomics research. However, the size of genomic datasets has grown exponentially over the last few years, leading to new challenges for traditional streaming algorithms. Modern petabyte-sized genomic datasets…
Juryon Paik, Junghyun Nam, Ung Mo Kim, Dongho Won
With the advances of wireless sensor networks, they yield massive volumes of disparate, dynamic and geographically-distributed and heterogeneous data. The data mining community has attempted to extract knowledge from the huge amount of data that they generate. However, previous mining work in WSNs has focused on…
Nirav Bhatt, Amit Thakkar, Othman Soufan
Stream data is the data that is generated continuously from the different data sources and ideally defined as the data that has no discrete beginning or end. Processing the stream data is a part of big data analytics that aims at querying the continuously arriving data and extracting meaningful information from the…
Glenda M. Yenni, Erica M. Christensen, Ellen K. Bledsoe, Sarah R. Supp + 3 more
Data management and publication are core components of the research process. An emerging challenge that has received limited attention in biology is managing, working with, and providing access to data under continual active collection. “Evolving data” present unique challenges in quality assurance and control, data…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Authors not listed
Machine learning models are transforming data-driven research across scientific disciplines, yet their deployment as accessible and reliable web services remains a significant challenge. We introduce the NERDD framework, a scalable, maintainable, and secure microservices platform designed to support the sustainable…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…