17 papers · ranked by Valyu relevance
Ameneh Zarei, Shahla Safari, Mahmood Ahmadi, Farhad Mardukhi
In this paper, a technology for massive data storage and computing named Hadoop is surveyed. Hadoop consists of heterogeneous computing devices like regular PCs abstracting away the details of parallel processing and developers can just concentrate on their computational problem. A Hadoop cluster is made of two parts…
Alexandros Gazis, Eleftheria Katsiri, Raffaele Bruno
This article introduces a novel middleware that utilizes cost-effective, low-power computing devices like Raspberry Pi to analyze data from wireless sensor networks (WSNs). It is designed for indoor settings like historical buildings and museums, tracking visitors and identifying points of interest. It serves as an…
Rajendra Purohit, K. R. Chowdhary, Sunıl Dutt Purohıt
—The parallel and distributed processing are becoming de facto industry standard, and a large part of the current research is targeted on how to make computing scalable and distributed, dynamically, without allocating the resources on permanent basis. The present article focuses on the study and performance of…
Angelos Dorotheos Chatzopoulos, Babis Andreou, Kakia Panagidi, Stathes Hadjiefthymiades
Modern logistics systems tend to generate continuous streams of data from sources such as GPS, IoT sensors, and logistics management systems. The aggregation, processing, and analysis of data have become vital for monitoring operations, optimizing efficiency, and responding quickly to decision making tasks. In this…
Ahmed Sharafeldeen, Mohammed Alrahmawy, Samir Elmougy
Counting number of triangles in the graph is considered a major task in many large-scale graph analytics problems such as clustering coefficient, transitivity ratio, trusses, etc. In recent years, MapReduce becomes one of the most popular and powerful frameworks for analyzing large-scale graphs in clusters of machines.…
Surendran Rajendran, Osamah Ibrahim Khalaf, Youseef Alotaibi, Saleh Alghamdi
In recent times, big data classification has become a hot research topic in various domains, such as healthcare, e-commerce, finance, etc. The inclusion of the feature selection process helps to improve the big data classification process and can be done by the use of metaheuristic optimization algorithms. This study…
Tarang Barasiya, Neeraj Bharti, Renu Gadhari, Avinash Bayaskar + 4 more
Large scale genome sequencing projects have produced huge datasets that pose challenges of high processing times especially for variant calling, a significant downstream analysis step. Efficient utilization of computational resources for accurate variant prediction in a timely manner is possible using Hadoop MapReduce…
Nithin Kavi
In this project, the goal was to use the Julia programming language and parallelization to write a fast map reduce algorithm to count word frequencies across large numbers of documents. We first implement the word frequency counter algorithm on a CPU using two processes with MPI. Then, we create another implementation…
R. Ramani, S. Edwin Raja, D. Dhinakaran, S. Jagan + 1 more
Recent trendy applications of Artificial Intelligence are Machine Learning (ML) algorithms, which have been extensively utilized for processes like pattern recognition, object classification, effective prediction of disease etc. However, ML techniques are reasonable solutions to computation methods and modeling…
Ziheng Wang, Atem Aguer, Amir Ziai
In this project we explore ways to dynamically load balance actors in a streaming framework. This is used to address input data skew that might lead to stragglers. We continuously monitor actors' input queue lengths for load, and redistribute inputs among reducers using consistent hashing if we detect stragglers. To…
Jun-Ha Lee, Hyuk-Yoon Kwon, Qingzhong Liu
In this study, we investigate large-scale digital forensic investigation on Apache Spark using a Windows registry. Because the Windows registry depends on the system on which it operates, the existing forensic methods on the Windows registry have been targeted on the Windows registry in a single system. However, it is…
Nicholas Kofi Akortia Hagan, John R. Talburt, Kris E. Anderson, Deasia Hagan
'Deasia Hagan'] Traditional data curation processes typically depend on human intervention. As data volume and variety grow exponentially, organizations are striving to increase efficiency of their data processes by automating manual processes and making them as unsupervised as possible. An additional challenge is to…
Meijing Li, Tianjie Chen, Keun Ho Ryu, Cheng Hao Jin
Semantic mining is always a challenge for big biomedical text data. Ontology has been widely proved and used to extract semantic information. However, the process of ontology-based semantic similarity calculation is so complex that it cannot measure the similarity for big text data. To solve this problem, we propose a…
Nicholas Kofi Akortia Hagan, John R. Talburt
Data volume has been one of the fast-growing assets of most real-world applications. This increases the rate of human errors such as duplication of records, misspellings, and erroneous transpositions, among other data quality issues. Entity Resolution is an ETL process that aims to resolve data inconsistencies by…
Dave Bunten, Jenna Tomkinson, Erik Serrano, Michael J. Lippincott + 4 more
High-content imaging (HCI) involves the automated acquisition and quantitative analysis of cell phenotypes from microscopy images. HCI has become a vital tool in biomedical research, supporting studies to characterize disease mechanisms and to identify new therapeutic agents. These studies often rely on high-throughput…
Tanveer Ahmad, Chengxin Ma, Zaid Al-Ars, H. Peter Hofstee
Current cluster scaled genomics data processing solutions rely on big data frameworks like Apache Spark, Hadoop and HDFS for data scheduling, processing and storage. These frameworks come with additional computation and memory overheads by default. It has been observed that scaling genomics dataset processing beyond 32…
Maxim Lippeveld, Daniel Peralta, Andrew Filby, Yvan Saeys
Due to high resolution and throughput of modern image cytometry platforms, morphologically profiling generated datasets poses a significant computational challenge. Here, we present Scalable Cytometry Image Processing (SCIP), an image processing software aimed at running on distributed high performance computing…