17 papers · ranked by Valyu relevance
Alexandros Gazis, Eleftheria Katsiri, Raffaele Bruno
This article introduces a novel middleware that utilizes cost-effective, low-power computing devices like Raspberry Pi to analyze data from wireless sensor networks (WSNs). It is designed for indoor settings like historical buildings and museums, tracking visitors and identifying points of interest. It serves as an…
Emad A Mohammed, Behrouz H Far, Christopher Naugler
The emergence of massive datasets in a clinical setting presents both challenges and opportunities in data storage and analysis. This so called “big data” challenges traditional analytic tools and will increasingly require novel solutions adapted from other fields. Advances in information and communication technology…
Wei Fang, V. S. Sheng, XueZhi Wen, Wubin Pan
In the atmospheric science, the scale of meteorological data is massive and growing rapidly. K-means is a fast and available cluster algorithm which has been used in many fields. However, for the large-scale meteorological data, the traditional K-means algorithm is not capable enough to satisfy the actual application…
Ahmed Sharafeldeen, Mohammed Alrahmawy, Samir Elmougy
Counting number of triangles in the graph is considered a major task in many large-scale graph analytics problems such as clustering coefficient, transitivity ratio, trusses, etc. In recent years, MapReduce becomes one of the most popular and powerful frameworks for analyzing large-scale graphs in clusters of machines.…
Suzanne J Matthews, Tiffani L Williams
Background MapReduce is a parallel framework that has been used effectively to design large-scale parallel applications for large computing clusters. In this paper, we evaluate the viability of the MapReduce framework for designing phylogenetic applications. The problem of interest is generating the all-to-all…
Ahmed Abdulhakim Al-Absi, Najeeb Abbas Al-Sammarraie, Wael Mohamed Shaher Yafooz, Dae-Ki Kang
MapReduce is the preferred cloud computing framework used in large data analysis and application processing. MapReduce frameworks currently in place suffer performance degradation due to the adoption of sequential processing approaches with little modification and thus exhibit underutilization of cloud resources. To…
Chang Sik Kim, Martyn D. Winn, Vipin Sachdeva, Kirk E. Jordan
Background De novo transcriptome assembly is an important technique for understanding gene expression in non-model organisms. Many de novo assemblers using the de Bruijn graph of a set of the RNA sequences rely on in-memory representation of this graph. However, current methods analyse the complete set of read-derived…
Muhammad Idris, Shujaat Hussain, Muhammad Hameed Siddiqi, Waseem Hassan + 3 more
'Waseem Hassan' 'Hafiz Syed Muhammad Bilal' 'Sungyoung Lee' 'Christophe Antoniewski'] Large quantities of data have been generated from multiple sources at exponential rates in the last few years. These data are generated at high velocity as real time and streaming data in variety of formats. These characteristics give…
Yunhong Gu, Robert L. Grossman
Cloud computing has demonstrated that processing very large datasets over commodity clusters can be done simply, given the right programming model and infrastructure. In this paper, we describe the design and implementation of the Sector storage cloud and the Sphere compute cloud. By contrast with the existing storage…
R. Ramani, S. Edwin Raja, D. Dhinakaran, S. Jagan + 1 more
Recent trendy applications of Artificial Intelligence are Machine Learning (ML) algorithms, which have been extensively utilized for processes like pattern recognition, object classification, effective prediction of disease etc. However, ML techniques are reasonable solutions to computation methods and modeling…
Yufei Gao, Yanjie Zhou, Bing Zhou, Lei Shi + 1 more
The healthcare industry has generated large amounts of data, and analyzing these has emerged as an important problem in recent years. The MapReduce programming model has been successfully used for big data analytics. However, data skew invariably occurs in big data analytics and seriously affects efficiency. To…
Jonas S Almeida, Alexander Grüneberg, Wolfgang Maass, Susana Vinga
Background The dramatic fall in the cost of genomic sequencing, and the increasing convenience of distributed cloud computing resources, positions the MapReduce coding pattern as a cornerstone of scalable bioinformatics algorithm development. In some cases an algorithm will find a natural distribution via use of map…
Weiyu Fu, Lixia Wang
Considering that in the process of job scheduling, the cluster load should be prebalanced rather than remedied when the load is seriously unbalanced; therefore, in this paper, the task scheduling flow of the Hadoop cluster is analyzed deeply. On the Hadoop platform, a self-dividing algorithm is proposed for load…
Shicai Wang, Ioannis Pandis, David Johnson, Ibrahim Emam + 3 more
'Florian Guitton' 'Axel Oehmichen' 'Yike Guo'] Background High-throughput molecular profiling data has been used to improve clinical decision making by stratifying subjects based on their molecular profiles. Unsupervised clustering algorithms can be used for stratification purposes. However, the current speed of the…
Jianfang Cao, Lichao Chen, Min Wang, Hao Shi + 1 more
Image classification uses computers to simulate human understanding and cognition of images by automatically categorizing images. This study proposes a faster image classification approach that parallelizes the traditional Adaboost-Backpropagation (BP) neural network using the MapReduce parallel programming model.…
Marco Capuccini, Martin Dahlö, Salman Toor, Ola Spjuth
MaRe comes as a thin layer on top of the RDD API , and it relies on Apache Spark to provide important features such as data locality, data ingestion, interactive processing, and fault tolerance. The implementation effort consists of (i) leveraging the RDD API to implement the MaRe primitives and (ii) handling data…
Martin Werner
This paper provides an abstract analysis of parallel processing strategies for spatial and spatio-temporal data. It isolates aspects such as data locality and computational locality as well as redundancy and locally sequential access as central elements of parallel algorithm design for spatial data. Furthermore, the…