16 papers · ranked by Valyu relevance
Sandeep Kumar, Sindhu Padakandla, L. Chandrashekar, Priyank Parihar + 2 more
'K. Gopinath' 'Shalabh Bhatnagar'] Hadoop MapReduce is a framework for distributed storage and processing of large datasets that is quite popular in big data analytics. It has various configuration parameters (knobs) which play an important role in deciding the performance i.e., the execution time of a given big data…
Ahmed Abdulhakim Al-Absi, Najeeb Abbas Al-Sammarraie, Wael Mohamed Shaher Yafooz, Dae-Ki Kang
MapReduce is the preferred cloud computing framework used in large data analysis and application processing. MapReduce frameworks currently in place suffer performance degradation due to the adoption of sequential processing approaches with little modification and thus exhibit underutilization of cloud resources. To…
Muhammad Idris, Shujaat Hussain, Muhammad Hameed Siddiqi, Waseem Hassan + 3 more
'Waseem Hassan' 'Hafiz Syed Muhammad Bilal' 'Sungyoung Lee' 'Christophe Antoniewski'] Large quantities of data have been generated from multiple sources at exponential rates in the last few years. These data are generated at high velocity as real time and streaming data in variety of formats. These characteristics give…
Yufei Gao, Yanjie Zhou, Bing Zhou, Lei Shi + 1 more
The healthcare industry has generated large amounts of data, and analyzing these has emerged as an important problem in recent years. The MapReduce programming model has been successfully used for big data analytics. However, data skew invariably occurs in big data analytics and seriously affects efficiency. To…
Jianfang Cao, Hongyan Cui, Hao Shi, Lijuan Jiao + 1 more
A back-propagation (BP) neural network can solve complicated random nonlinear mapping problems; therefore, it can be applied to a wide range of problems. However, as the sample size increases, the time required to train BP neural networks becomes lengthy. Moreover, the classification accuracy decreases as well. To…
Vaneet Aggarwal, Tian Lan, Suresh Subramaniam, Maotong Xu
—MapReduce is the most popular big-data computation framework, motivating many research topics. A MapReduce job consists of two successive phases, i.e., map phase and reduce phase. Each phase can be divided into multiple tasks. A reduce task can only start when all the map tasks finish processing. A job is successfully…
Zhuo Wang, Longlong Tian, Dianjie Guo, Xiaoming Jiang
When dealing with massive data sorting, we usually use Hadoop which is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. A common approach in implement of big data sorting is to use shuffle and sort phase in MapReduce based on Hadoop.…
Güngör Yıldırım, İbrahim Rıza Hallaç, Galip Aydın, Yetkin Tatar
—Hadoop is a popular MapReduce framework for developing parallel applications in distributed environments. Several advantages of MapReduce such as programming ease and ability to use commodity hardware make the applicability of soft computing methods for parallel and distributed systems easier than before. In this…
Deepak Narayanan, Fiodar Kazhamiaka, Firas Abuzaid, Peter Kraft + 1 more
'Matei Zaharia'] Resource allocation problems in many computer systems can be formulated as mathematical optimization problems. However, finding exact solutions to these problems using off-the-shelf solvers in an online setting is often intractable for "hyper-scale" system sizes with tight SLAs, leading system…
Mian Lu, Lei Zhang, Huynh Phung Huynh, Zhongliang Ong + 4 more
'Bingsheng He' 'Rick Siow Mong Goh' 'Richard Huynh'] Abstract—With the ease-of-programming, flexibility and yet efficiency, MapReduce has become one of the most popular frameworks for building big-data applications. MapReduce was originally designed for distributed-computing, and has been extended to various…
Qinghua Lu, Shanshan Li, Weishan Zhang, Lei Zhang
Big data analytics (BDA) applications are a new category of software applications that process large amounts of data using scalable parallel processing infrastructure to obtain hidden value. Hadoop is the most mature open-source big data analytics framework, which implements the MapReduce programming model to process…
Jamie Alnasir, Hugh P. Shanahan
The paper reviews the use of the Hadoop platform in Structural Bioinformatics applications. Specifically, we review a number of implementations using Hadoop of high-throughput analyses, e.g. ligand-protein docking and structural alignment, and their scalability in comparison with other batch schedulers and MPI. We find…
Weiyu Fu, Lixia Wang
Considering that in the process of job scheduling, the cluster load should be prebalanced rather than remedied when the load is seriously unbalanced; therefore, in this paper, the task scheduling flow of the Hadoop cluster is analyzed deeply. On the Hadoop platform, a self-dividing algorithm is proposed for load…
Abbas Kazemipour, Behtash Babadi, Min Wu, Kaspar Podgorski + 1 more
We consider the problem of optimizing general convex objective functions with nonnegativity constraints. Using the Karush-Kuhn-Tucker (KKT) conditions for the nonnegativity constraints we will derive fast multiplicative update rules for several problems of interest in signal processing, including non-negative…
Mary Pitman, David Hahn, Gary Tresadern, David Mobley
Drug discovery is accelerated with computational methods such as alchemical simulations to estimate ligand affinities. In particular, relative binding free energy (RBFE) simulations are beneficial for lead optimization. To use RBFE simulations to compare prospective ligands in silico, researchers first plan the…
Murat Okatan
The firing rate of hippocampal place cells depends on the spatial position of the organism in an environment. This position dependence is often quantified by constructing spike-in-location and time-in-location histograms, the ratio of which yields a firing rate map. The purpose of this study is to present a new method…