15 papers · ranked by Valyu relevance
Khoat Than
In this work, we develop a theoretical framework that elucidates the role of normalization through the lens of capacity control. We prove that an unnormalized DNN can exhibit exponentially large Lipschitz constants with respect to either its parameters or inputs, implying excessive functional capacity and potential…
Jiacheng Sun, Xiangyong Cao, Hanwen Liang, Weiran Huang + 2 more
'Zhenguo Li'] In recent years, a variety of normalization methods have been proposed to help train neural networks, such as batch normalization (BN), layer normalization (LN), weight normalization (WN), group normalization (GN), etc. However, mathematical tools to analyze all these normalization methods are lacking. In…
Mengye Ren, Renjie Liao, Raquel Urtasun, Fabian H. Sinz + 1 more
'Richard S. Zemel'] Normalization techniques have only recently begun to be exploited in supervised learning tasks. Batch normalization exploits mini-batch statistics to normalize the activations. This was shown to speed up training and result in better models. However its success has been very limited when dealing…
Max Robinson
A numerical data matrix may be seen simply as a means of organizing observations of a system into rows in one manner (e.g., by measured "object"), and into columns by another (e.g., by measured "variable") so that the observations can be analyzed with mathematical tools. As a mathematical object, however, a matrix…
Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka
Inspired by BatchNorm, there has been an explosion of normalization layers in deep learning. Recent works have identified a multitude of beneficial properties in BatchNorm to explain its success. However, given the pursuit of alternative normalization layers, these properties need to be generalized so that any given…
Jiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun + 1 more
Normalization layers are ubiquitous in modern neural networks and have long been considered essential. This work demonstrates that Transformers without normalization can achieve the same or better performance using a remarkably simple technique. We introduce Dynamic Tanh (DyT), an element-wise operation DyT(x) =…
Lucas Deecke, Iain Murray, Hakan Bilen
Normalization methods are a central building block in the deep learning toolbox. They accelerate and stabilize training, while decreasing the dependence on manually tuned learning rate schedules. When learning from multi-modal distributions, the effectiveness of batch normalization (BN), arguably the most prominent…
Alexander Shekhovtsov, Boris Flach
We address the problem of estimating statistics of hidden units in a neural network using a method of analytic moment propagation. These statistics are useful for approximate whitening of the inputs in front of saturating non-linearities such as a sigmoid function. This is important for initialization of training and…
Amir Joudaki, Hadi Daneshmand, Francis Bach
In this paper, we explore the structure of the penultimate Gram matrix in deep neural networks, which contains the pairwise inner products of outputs corresponding to a batch of inputs. In several architectures it has been observed that this Gram matrix becomes degenerate with depth at initialization, which…
Massimiliano Esposito, Nader Ganaba
As the deep neural networks are being applied to complex tasks, the size of the networks and architecture increases and their topology becomes more complicated too. At the same time, training becomes slow and at some instances inefficient. This motivated the introduction of various normalization techniques such as…
Rudrasis Chakraborty
Many measurements in computer vision and machine learning manifest as non-Euclidean data samples. Several researchers recently extended a number of deep neural network architectures for manifold valued data samples. Researchers have proposed models for manifold valued spatial data which are common in medical image…
Daniel Eftekhari, Vardan Papyan
The normal distribution plays a central role in information theory – it is at the same time the best-case signal and worst-case noise distribution, has the greatest representational capacity of any distribution, and offers an equivalence between uncorrelatedness and independence for joint distributions. Accounting for…
Róbert Rajkó
This short communication paper is briefly dealing with greater or lesser misused normalization in self-modeling/multivariate curve resolution (S/MCR) practice. The importance of the correct use of the ODE (ordinary differential equation) solvers and apt kinetic illustrations are elucidated. The new terms, external and…
Sébastien Herbreteau, Emmanuel Moebel, Charles Kervrann
In many information processing systems, it may be desirable to ensure that any change of the input, whether by shifting or scaling, results in a corresponding change in the system response. While deep neural networks are gradually replacing all traditional automatic processing methods, they surprisingly do not…
Modis, Theodore
Use is made of rigorous definitions for the terms normal, natural, and harmonic to reveal a number of unfamiliar aspects about them. The Gaussian distribution is not sufficient to determine who is normal, and fluctuations above or below a natural-growth curve may or may not be natural. A recipe for harmonically…