14 papers · ranked by Valyu relevance
Baturalp Buyukates, Emre Özfatura, Şennur Ulukuş, Denız Gündüz
Distributed implementations are crucial in speeding up large scale machine learning applications. Distributed gradient descent (GD) is widely employed to parallelize the learning task by distributing the dataset across multiple workers. A significant performance bottleneck for the per-iteration completion time in…
Qi Wang, Ying Cui, Chenglin Li, Junni Zou + 1 more
—Existing gradient coding schemes introduce identical redundancy across the coordinates of gradients and hence cannot fully utilize the computation results from partial stragglers. This motivates the introduction of diverse redundancies across the coordinates of gradients. This paper considers a distributed computation…
Yuxin Jiang, Wenqin Zhang, Lele Wang
Gradient coding is a distributed computing technique aiming to provide robustness against slow or non-responsive computing nodes, known as stragglers, while balancing the computational load for responsive computing nodes. Among existing gradient codes, a construction based on combinatorial designs, called BIBD gradient…
Luis Maßny, Christoph Hofmeister, Maximilian Egger, Rawad Bitar + 1 more
'Antonia Wachter-Zeh'] Abstract—We consider distributed learning in the presence of slow and unresponsive worker nodes, referred to as stragglers. In order to mitigate the effect of stragglers, gradient coding redundantly assigns partial computations to the worker such that the overall result can be recovered from only…
Muhammet Balcılar, Bharath Bhushan Damodaran, Karam Naser, Franck Galpin + 1 more
'Franck Galpin' 'Pierre Hellier'] End-to-end image/video codecs are getting competitive compared to traditional compression techniques that have been developed through decades of manual engineering efforts. These trainable codecs have many advantages over traditional techniques such as easy adaptation on perceptual…
Jannis Clausius, Marvin Geiselhart, Stephan ten Brink
—For improving short-length codes, we demonstrate that classic decoders can also be used with real-valued, neural encoders, i.e., deep-learning based "codeword" sequence generators. Here, the classical decoder can be a valuable tool to gain insights into these neural codes and shed light on weaknesses. Specifically…
Louis-Adrien Dufrène, Quentin Lampin, Guillaume Larue
—This study investigates the problem of learning linear block codes optimized for Belief-Propagation decoders significantly improving performance compared to the state-ofthe-art. Our previous research is extended with an enhanced system design that facilitates a more effective learning process for the parity check…
Tadashi Wadayama, Lantian Wei
—This paper presents the Gradient Flow (GF) decoding for LDPC codes. GF decoding, a continuous-time methodology based on gradient flow, employs a potential energy function associated with bipolar codewords of LDPC codes. The decoding process of the GF decoding is concisely defined by an ordinary differential equation…
Christopher Fifty, Ronald G. Junkins, D. Duan, Aniketh Iger + 4 more
'Ehsan Amid' 'Sebastian Thrun' 'Christopher Ré'] Vector Quantized Variational AutoEncoders (VQ-VAEs) are designed to compress a continuous input to a discrete latent space and reconstruct it with minimal distortion. They operate by maintaining a set of vectors—often referred to as the codebook—and quantizing each…
Muhammet Balcılar, Bharath Bhushan Damodaran, Karam Naser, Franck Galpin + 1 more
'Franck Galpin' 'Pierre Hellier'] Abstract—End-to-end image and video codecs are becoming increasingly competitive, compared to traditional compression techniques that have been developed through decades of manual engineering efforts. These trainable codecs have many advantages over traditional techniques, such as…
Tony Shaska
We introduce the Graded Transformer framework, embedding algebraic inductive biases via grading transformations on vector spaces. Extending Graded Neural Networks (GNNs), we propose the Linearly Graded Transformer (LGT) and Exponentially Graded Transformer (EGT), which apply parameterized scaling—via fixed or learnable…
Gergely Flamich, Stratis Markou, Jose Miguel Hernandez Lobato
Relative entropy coding (REC) algorithms encode a sample from a target distribution Q using a proposal distribution P using as few bits as possible. Unlike entropy coding, REC does not assume discrete distributions or require quantisation. As such, it can be naturally integrated into communication pipelines such as…
Hyunmin Cho, Jaejun Yoo, Kyong Hwan Jin
We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveals that sinusoidal activations induce a harmonic line spectrum, providing a spectral account of how recurrent unrolling enriches the effective spectral support. We…
Deval Shah, Zi Yu Xue, Tor M. Aamodt
Deep neural networks are used for a wide range of regression problems. However, there exists a significant gap in accuracy between specialized approaches and generic direct regression in which a network is trained by minimizing the squared or absolute error of output labels. Prior work has shown that solving a…