16 papers · ranked by Valyu relevance
Rajendra Purohit, K. R. Chowdhary, Sunıl Dutt Purohıt
—The parallel and distributed processing are becoming de facto industry standard, and a large part of the current research is targeted on how to make computing scalable and distributed, dynamically, without allocating the resources on permanent basis. The present article focuses on the study and performance of…
Ameneh Zarei, Shahla Safari, Mahmood Ahmadi, Farhad Mardukhi
In this paper, a technology for massive data storage and computing named Hadoop is surveyed. Hadoop consists of heterogeneous computing devices like regular PCs abstracting away the details of parallel processing and developers can just concentrate on their computational problem. A Hadoop cluster is made of two parts…
Philippe Besse, Nathalie Villa‐Vialaneix
Résumé : Cet article suppose acquises les compétences et expertises d'un statisticien en apprentissage non supervisé (NMF, k-means, svd) et supervisé (régression, cart, random forest). Quelles compétences et savoir faire ce statisticien doit-il acquérir pour passer à l'échelle "Volume" des grandes masses de données ?…
Sherif Sakr, Anna Liu, Ayman G. Fayoumi
In the last two decades, the continuous increase of computational power has produced an overwhelming flow of data which has called for a paradigm shift in the computing architecture and large scale data processing mechanisms. MapReduce is a simple and powerful programming model that enables easy development of scalable…
Colin Barrett, Christos Kotselidis, Mikel Luján
The explosion of Big Data was followed by the proliferation of numerous complex parallel software stacks whose aim is to tackle the challenges of data deluge. A drawback of a such multi-layered hierarchical deployment is the inability to maintain and delegate vital semantic information between layers in the stack.…
Yanfeng Zhang, Shimin Chen, Qiang Wang, Ge Yu
In this paper, we propose i2MapReduce, a novel incremental processing extension to MapReduce, the most widely used framework for mining big data. Compared with the state-of-the-art work on Incoop, i2MapReduce (i) performs key-value pair level incremental processing rather than task level re-computation, (ii) supports…
Matthew Felice Pace
The MapReduce framework has been generating a lot of interest in a wide range of areas. It has been widely adopted in industry and has been used to solve a number of non-trivial problems in academia. Putting MapReduce on strong theoretical foundations is crucial in understanding its capabilities. This work links…
Philip Derbeko, Shlomi Dolev, Ehud Gudes, Shantanu Sharma
MapReduce is a programming system for distributed processing large-scale data in an efficient and fault tolerant manner on a private, public, or hybrid cloud. MapReduce is extensively used daily around the world as an efficient distributed computation tool for a large class of problems, e.g., search, clustering, log…
Woo-Hyun Lee, Hee-Gook Jun, Hyoung-Joo Kim
While advanced analysis of large dataset is in high demand, data sizes have surpassed capabilities of conventional software and hardware. Hadoop framework distributes large datasets over multiple commodity servers and performs parallel computations. We discuss the I/O bottlenecks of Hadoop framework and propose methods…
Foto Afrati, Shlomi Dolev, Shantanu Sharma, Jeffrey D. Ullman
—MapReduce has proven to be one of the most useful paradigms in the revolution of distributed computing, where cloud services and cluster computing become the standard venue for computing. The federation of cloud and big data activities is the next challenge where MapReduce should be modified to avoid (big) data…
Nandan Mirajkar, Sandeep Bhujbal, Aaradhana Deshmukh
Applications like Yahoo, Facebook, Twitter have huge data which has to be stored and retrieved as per client access. This huge data storage requires huge database leading to increase in physical storage and becomes complex for analysis required in business growth. This storage capacity can be reduced and distributed…
Angelos Dorotheos Chatzopoulos, Babis Andreou, Kakia Panagidi, Stathes Hadjiefthymiades
Modern logistics systems tend to generate continuous streams of data from sources such as GPS, IoT sensors, and logistics management systems. The aggregation, processing, and analysis of data have become vital for monitoring operations, optimizing efficiency, and responding quickly to decision making tasks. In this…
Zhimin Yao, Blesson Varghese, Andrew Rau‐Chaplin
Monte Carlo simulations employed for the analysis of portfolios of catastrophic risk process large volumes of data. Often times these simulations are not performed in real-time scenarios as they are slow and consume large data. Such simulations can benefit from a framework that exploits parallelism for addressing the…
Nikzad Babaii Rizvandi, Albert Y. Zomaya, Ali Javadzadeh Boloori, Javid Taheri
'Javid Taheri'] Abstract—In this paper, we propose an analytical method to model the dependency between configuration parameters and total execution time of Map-Reduce applications. Our approach has three key phases: profiling, modeling, and prediction. In profiling, an application is run several times with different…
Junhao Li, Hang Zhang
MapReduce and its variants have significantly simplified and accelerated the process of developing parallel programs. However, most MapReduce implementations focus on data-intensive tasks while many real-world tasks are compute intensive and their data can fit distributedly into the memory. For these tasks, the speed…
Liya Fan, Bo Gao, Fa Zhang, Xiaogang Li
—The efficiency of MapReduce is closely related to its load balance. Existing works on MapReduce load balance focus on coarse-grained scheduling. This study concerns finegrained scheduling on MapReduce operations, with each operation representing one invocation of the Map or Reduce function. By default, MapReduce…