Search · four archives
Search · four archives
15 papers · ranked by Valyu relevance
Madhu S. Advani, Andrew M. Saxe, Haim Sompolinsky
We perform an analysis of the average generalization dynamics of large neural networks trained using gradient descent. We study the practically-relevant “high-dimensional” regime where the number of free parameters in the network is on the order of or even larger than the number of examples in the dataset. Using random…
Christopher W. Bartlett, Jamie Bossenbroek, Yukie Ueyama, Patricia McCallinhart + 5 more
Early stopping is an extremely common tool to minimize overfitting, which would otherwise be a cause of poor generalization of the model to novel data. However, early stopping is a heuristic that, while effective, primarily relies on ad hoc parameters and metrics. Optimizing when to stop remains a challenge. In this…
Xinkai Sun, Sanguo Zhang, Shuangge Ma, Sotiris Kotsiantis
In the classification task, label noise has a significant impact on models’ performance, primarily manifested in the disruption of prediction consistency, thereby reducing the classification accuracy. This work introduces a novel prediction consistency regularization that mitigates the impact of label noise on neural…
Abdulaziz Albahr, Marwan Albahar, Mohammed Thanoon, Muhammad Binsawad
'Muhammad Binsawad'] Heart diseases are characterized as heterogeneous diseases comprising multiple subtypes. Early diagnosis and prognosis of heart disease are essential to facilitate the clinical management of patients. In this research, a new computational model for predicting early heart disease is proposed. The…
Yuyang Gao, Giorgio A. Ascoli, Liang Zhao
In an attempt to test the influence of BEAN regularization on the generalizability of DNNs in the scenarios where the training samples are extremely limited, we conducted a few-shot learning from scratch task, i.e., without the help of any additional side tasks and pre-trained models (Kimura et al., ). Notice that in…
Alexander Effland, Erich Kobler, Karl Kunisch, Thomas Pock
We investigate a well-known phenomenon of variational approaches in image processing, where typically the best image quality is achieved when the gradient flow process is stopped before converging to a stationary point. This paradox originates from a tradeoff between optimization and modeling errors of the underlying…
Nassim Sohaee
In Machine Learning, prediction quality is usually measured using different techniques and evaluation methods. In the regression models, the goal is to minimize the distance between the actual and predicted value. This error evaluation technique lacks a detailed evaluation of the type of errors that occur on specific…
Abdulkadir Canatar, Blake Bordelon, Cengiz Pehlevan
A theoretical understanding of generalization remains an open problem for many machine learning models, including deep networks where overparameterization leads to better performance, contradicting the conventional wisdom from classical statistics. Here, we investigate generalization error for kernel regression, which…
Gauri Jagatap, Ameya Joshi, Animesh Basak Chowdhury, Siddharth Garg + 1 more
'Chinmay Hegde'] In this paper we propose a new family of algorithms, ATENT, for training adversarially robust deep neural networks. We formulate a new loss function that is equipped with an additional entropic regularization. Our loss function considers the contribution of adversarial samples that are drawn from a…
Yuhua Fan, Ilkka Launonen, Mikko J Sillanpää, Patrik Waldmann
High-dimensional genomic datasets contain complex patterns shaped by substantial biological noise, which pose major challenges for predictive modeling in genetics and breeding. Residual neural networks (ResNets) provide a powerful framework for capturing nonlinear genomic effects, but often overfit in settings where…
Omisa Jinsi, Margaret M. Henderson, Michael J. Tarr, Sathishkumar V. E.
'Sathishkumar V. E.'] Humans are born with very low contrast sensitivity, meaning that inputs to the infant visual system are both blurry and low contrast. Is this solely a byproduct of maturational processes or is there a functional advantage for beginning life with poor visual acuity? We addressed the impact of poor…
Jianwei Liu, Shuang Cheng Li, Xionglin Luo
Support vector machine is an effective classification and regression method that uses machine learning theory to maximize the predictive accuracy while avoiding overfitting of data. L2 regularization has been commonly used. If the training dataset contains many noise variables, L1 regularization SVM will provide a…
Johannes Schwab, Stephan Antholzer, Markus Haltmeier
Deep learning and (deep) neural networks are emerging tools to address inverse problems and image reconstruction tasks. Despite outstanding performance, the mathematical analysis for solving inverse problems by neural networks is mostly missing. In this paper, we introduce and rigorously analyze families of deep…
Weifeng Liu, Yang Li, Xu Lin, Dacheng Tao + 2 more
'Kewei Chen'] Co-training is a major multi-view learning paradigm that alternately trains two classifiers on two distinct views and maximizes the mutual agreement on the two-view unlabeled data. Traditional co-training algorithms usually train a learner on each view separately and then force the learners to be…
Victor Vergnieux, Rufin Vogels
Animals of several species, including primates, learn the statistical regularities of their environment. In particular, they learn the temporal regularities that occur in streams of visual images. Previous human neuroimaging studies reported discrepant effects of such statistical learning, ranging from stronger…