20 papers · ranked by Valyu relevance
Alexandros Gazis, Eleftheria Katsiri, Raffaele Bruno
This article introduces a novel middleware that utilizes cost-effective, low-power computing devices like Raspberry Pi to analyze data from wireless sensor networks (WSNs). It is designed for indoor settings like historical buildings and museums, tracking visitors and identifying points of interest. It serves as an…
Emad A Mohammed, Behrouz H Far, Christopher Naugler
The emergence of massive datasets in a clinical setting presents both challenges and opportunities in data storage and analysis. This so called “big data” challenges traditional analytic tools and will increasingly require novel solutions adapted from other fields. Advances in information and communication technology…
Rajendra Purohit, K. R. Chowdhary, Sunıl Dutt Purohıt
—The parallel and distributed processing are becoming de facto industry standard, and a large part of the current research is targeted on how to make computing scalable and distributed, dynamically, without allocating the resources on permanent basis. The present article focuses on the study and performance of…
Ameneh Zarei, Shahla Safari, Mahmood Ahmadi, Farhad Mardukhi
In this paper, a technology for massive data storage and computing named Hadoop is surveyed. Hadoop consists of heterogeneous computing devices like regular PCs abstracting away the details of parallel processing and developers can just concentrate on their computational problem. A Hadoop cluster is made of two parts…
Sherif Sakr, Anna Liu, Ayman G. Fayoumi
In the last two decades, the continuous increase of computational power has produced an overwhelming flow of data which has called for a paradigm shift in the computing architecture and large scale data processing mechanisms. MapReduce is a simple and powerful programming model that enables easy development of scalable…
Colin Barrett, Christos Kotselidis, Mikel Luján
The explosion of Big Data was followed by the proliferation of numerous complex parallel software stacks whose aim is to tackle the challenges of data deluge. A drawback of a such multi-layered hierarchical deployment is the inability to maintain and delegate vital semantic information between layers in the stack.…
Wei Fang, V. S. Sheng, XueZhi Wen, Wubin Pan
In the atmospheric science, the scale of meteorological data is massive and growing rapidly. K-means is a fast and available cluster algorithm which has been used in many fields. However, for the large-scale meteorological data, the traditional K-means algorithm is not capable enough to satisfy the actual application…
Woo-Hyun Lee, Hee-Gook Jun, Hyoung-Joo Kim
While advanced analysis of large dataset is in high demand, data sizes have surpassed capabilities of conventional software and hardware. Hadoop framework distributes large datasets over multiple commodity servers and performs parallel computations. We discuss the I/O bottlenecks of Hadoop framework and propose methods…
Foto Afrati, Shlomi Dolev, Shantanu Sharma, Jeffrey D. Ullman
—MapReduce has proven to be one of the most useful paradigms in the revolution of distributed computing, where cloud services and cluster computing become the standard venue for computing. The federation of cloud and big data activities is the next challenge where MapReduce should be modified to avoid (big) data…
Nandan Mirajkar, Sandeep Bhujbal, Aaradhana Deshmukh
Applications like Yahoo, Facebook, Twitter have huge data which has to be stored and retrieved as per client access. This huge data storage requires huge database leading to increase in physical storage and becomes complex for analysis required in business growth. This storage capacity can be reduced and distributed…
Suzanne J Matthews, Tiffani L Williams
Background MapReduce is a parallel framework that has been used effectively to design large-scale parallel applications for large computing clusters. In this paper, we evaluate the viability of the MapReduce framework for designing phylogenetic applications. The problem of interest is generating the all-to-all…
Muhammad Idris, Shujaat Hussain, Muhammad Hameed Siddiqi, Waseem Hassan + 3 more
'Waseem Hassan' 'Hafiz Syed Muhammad Bilal' 'Sungyoung Lee' 'Christophe Antoniewski'] Large quantities of data have been generated from multiple sources at exponential rates in the last few years. These data are generated at high velocity as real time and streaming data in variety of formats. These characteristics give…
Chang Sik Kim, Martyn D. Winn, Vipin Sachdeva, Kirk E. Jordan
De novo transcriptome assembly is an important technique for understanding gene expression in non-model organisms. Many de novo assemblers using the de Bruijn graph of a set of the RNA sequences rely on in-memory representation of this graph. However, current methods analyse the complete set of read-derived k-mer…
Tarang Barasiya, Neeraj Bharti, Renu Gadhari, Avinash Bayaskar + 4 more
Large scale genome sequencing projects have produced huge datasets that pose challenges of high processing times especially for variant calling, a significant downstream analysis step. Efficient utilization of computational resources for accurate variant prediction in a timely manner is possible using Hadoop MapReduce…
Jamie Alnasir, Hugh P. Shanahan
The paper reviews the use of the Hadoop platform in Structural Bioinformatics applications. Specifically, we review a number of implementations using Hadoop of high-throughput analyses, e.g. ligand-protein docking and structural alignment, and their scalability in comparison with other batch schedulers and MPI. We find…
Weiyu Fu, Lixia Wang
Considering that in the process of job scheduling, the cluster load should be prebalanced rather than remedied when the load is seriously unbalanced; therefore, in this paper, the task scheduling flow of the Hadoop cluster is analyzed deeply. On the Hadoop platform, a self-dividing algorithm is proposed for load…
Martin Werner
This paper provides an abstract analysis of parallel processing strategies for spatial and spatio-temporal data. It isolates aspects such as data locality and computational locality as well as redundancy and locally sequential access as central elements of parallel algorithm design for spatial data. Furthermore, the…
Francesco Versaci, Luca Pireddu, Gianluigi Zanetti
The adoption of Big Data technologies can potentially boost the scalability of data-driven biology and health workflows by orders of magnitude. Consider, for instance, that technologies in the Hadoop ecosystem have been successfully used in data-driven industry to scale their processes to levels much larger than any…
Miroslav Kratochvíl, Oliver Hunewald, Laurent Heirendt, Vasco Verissimo + 5 more
The amount of data generated in large clinical and phenotyping studies that use single-cell cytometry is constantly growing. Recent technological advances allow to easily generate data with hundreds of millions of single-cell data points with more than 40 parameters, originating from thousands of individual samples.…
Mika Sarkin Jain, Cecilia Dominguez Conde, Krzysztof Polanski, Xi Chen + 7 more
Multimodal data is rapidly growing in many fields of science and engineering, including single-cell biology. We introduce MultiMAP, an approach for the dimensionality reduction and integration of multiple datasets. MultiMAP recovers a single manifold on which all of the data resides and then projects the data into a…