Search · four archives
Search · four archives
19 papers · ranked by Valyu relevance
Kangwook Lee, Maximilian Lam, Ramtin Pedarsani, Dimitris Papailiopoulos + 1 more
'Dimitris Papailiopoulos' 'Kannan Ramchandran'] Codes are widely used in many engineering applications to offer robustness against noise. In large-scale systems there are several types of noise that can affect the performance of distributed machine learning algorithms – straggler nodes, system failures, or…
Yongcheng Yang, Yifei Huang, Xiaohuan Qin, Shenglian Lu + 3 more
'Yanlin Geng' 'Youlong Wu' 'Ling Liu'] Coded distributed computing (CDC) is a powerful approach to reduce the communication overhead in distributed computing frameworks by utilizing coding techniques. In this paper, we focus on the CDC problem in $(H,L)$-combination networks, where H APs act as intermediate pivots and…
Songze Li, Mohammad Ali Maddah-Ali, A. Salman Avestimehr
More specifically, a general distributed computing framework, motivated by commonly used structures like MapReduce, is considered, where the overall computation is decomposed into computing a set of "Map" and "Reduce" functions distributedly across multiple computing nodes. A coded scheme, named "Coded Distributed…
Kai Wan, Mingyue Ji, Giuseppe Caire
—This paper considers the MapReduce-like coded distributed computing framework originally proposed by Li et al., which uses coding techniques when distributed computing servers exchange their computed intermediate values, in order to reduce the overall traffic load. Their original model servers are connected via an…
Yingjie Cheng, Gaojun Luo, Xiwang Cao, Martianus Frederic Ezerman + 1 more
'San Ling'] A coded distributed computing (CDC) system aims to reduce the communication load in the MapReduce framework. Such a system has K nodes, N input files, and Q Reduce functions. Each input file is mapped by r nodes and each Reduce function is computed by s nodes. The objective is to achieve the maximum…
Qicheng Zeng, Zhaojun Nan, Sheng Zhou, T. Aaron Gulliver
Coded computing is recognized as a promising solution to address the privacy leakage problem and the straggling effect in distributed computing. This technique leverages coding theory to recover computation tasks using results from a subset of workers. In this paper, we propose the adaptive privacy-preserving coded…
Yingjie Cheng, Gaojun Luo, Xiwang Cao, Martianus Frederic Ezerman + 1 more
'San Ling'] Coded distributed computing (CDC) was introduced to greatly reduce the communication load for MapReduce computing systems. Such a system has K nodes, N input files, and Q Reduce functions. Each input file is mapped by r nodes and each Reduce function is computed by s nodes. The architecture must allow for…
Songze Li, Sucha Supittayapornpong, Mohammad Ali Maddah-Ali, A. Salman Avestimehr
'A. Salman Avestimehr'] Abstract—We focus on sorting, which is the building block of many machine learning algorithms, and propose a novel distributed sorting algorithm, named CodedTeraSort, which substantially improves the execution time of the TeraSort benchmark in Hadoop MapReduce. The key idea of CodedTeraSort is…
Jer Shyuan Ng, Wei Yang Bryan Lim, Nguyen Cong Luong, Zehui Xiong + 4 more
'Alia Asheralieva' 'Dusit Niyato' 'Cyril Leung' 'Chunyan Miao'] Abstract—Distributed computing has become a common approach for large-scale computation of tasks due to benefits such as high reliability, scalability, computation speed, and costeffectiveness. However, distributed computing faces critical issues related…
Bin Fan, Bin Tang, Zhihao Qu, Baoliu Ye + 1 more
In wireless distributed computing systems, worker nodes connect to a master node wirelessly and perform large-scale computational tasks that are parallelized across them. However, the common phenomenon of straggling (i.e., worker nodes often experience unpredictable slowdown during computation and communication) and…
Ming Xiao, Mikael Skoglund, H. Vincent Poor, Onur Günlü + 2 more
'Rafael F. Schaefer' 'Holger Boche'] This article aims to give a comprehensive and rigorous review of the principles and recent development of coding for large-scale distributed machine learning (DML). With increasing data volumes and the pervasive deployment of sensors and computing machines, machine learning has…
Derya Malak, Mohammad Reza Deylam Salehi, Berksan Serbetci, Petros Elia + 2 more
'Petros Elia' 'Chintha Tellambura' 'Jun Chen'] The work here studies the communication cost for a multi-server multi-task distributed computation framework, as well as for a broad class of functions and data statistics. Considering the framework where a user seeks the computation of multiple complex (conceivably…
Emre Ozfatura, Sennur Ulukus, Deniz Gündüz
When gradient descent (GD) is scaled to many parallel workers for large-scale machine learning applications, its per-iteration computation time is limited by straggling workers. Straggling workers can be tolerated by assigning redundant computations and/or coding across data and computations, but in most existing…
Prasad Krishnan, Lakshmi Natarajan, V. Lalitha, Siu-Wai Ho + 2 more
'Lawrence Ong' 'Kenneth Shum'] The problem of data exchange between multiple nodes with storage and communication capabilities models several current multi-user communication problems like Coded Caching, Data Shuffling, Coded Computing, etc. The goal in such problems is to design communication schemes which accomplish…
Jia Lu, Ryan Tsoi, Nan Luo, Yuanchi Ha + 8 more
Dynamical systems often generate distinct outputs according to different initial conditions, and one can infer the corresponding input configuration given an output. This property captures the essence of information encoding and decoding. Here, we demonstrate the use of self-organized patterns, combined with machine…
Timothy C. Haas
Models of political-ecological systems can inform policies for managing ecosystems that contain endangered species. One way to increase the credibility of these models is to subject them to a rigorous suite of data-based statistical assessments. Doing so involves statistically estimating the model’s parameters…
Pierre Carrier, Bill Long, Richard Walsh, Jef Dawson + 4 more
High Performance Computing (HPC) Best Practice offers opportunities to implement lessons learned in areas such as computational chemistry and physics in genomics workflows, specifically Next-Generation Sequencing (NGS) workflows. In this study we will briefly describe how distributed-memory parallelism can be an…
Cláudia Brito, Pedro Ferreira, João Paulo
Breakthroughs in sequencing technologies led to an exponential growth of genomic data, providing unprecedented biological in-sights and new therapeutic applications. However, analyzing such large amounts of sensitive data raises key concerns regarding data privacy, specifically when the information is outsourced to…
Authors not listed
Machine learning models are transforming data-driven research across scientific disciplines, yet their deployment as accessible and reliable web services remains a significant challenge. We introduce the NERDD framework, a scalable, maintainable, and secure microservices platform designed to support the sustainable…