Search · four archives
Search · four archives
23 papers · ranked by Valyu relevance
Yunji Yang, Yonggi Hong, Jaehyun Park, Leopoldo Angrisani
In this paper, efficient gradient updating strategies are developed for the federated learning when distributed clients are connected to the server via a wireless backhaul link. Specifically, a common convolutional neural network (CNN) module is shared for all the distributed clients and it is trained through the…
Yash Ganpat Sawant
Adaptive rank allocation for LoRA, allocating more parameters to important layers and fewer to unimportant ones, consistently improves efficiency under supervised fine-tuning (SFT). We investigate whether this success transfers to reinforcement learning, specifically Group Relative Policy Optimization (GRPO). Using…
A. R. Flores, Rodrigo C. de Lamare
Rate splitting (RS) systems can better deal with imperfect channel state information at the transmitter (CSIT) than conventional approaches. However, this requires an appropriate power allocation that often has a high computational complexity, which might be inadequate for practical and large systems. To this end…
Ali Kavis, Kfir Y. Levy, Volkan Cevher
In this paper, we propose a new, simplified high probability analysis of AdaGrad for smooth, non-convex problems. More specifically, we focus on a particular accelerated gradient (AGD) template (Lan , 2020), through which we recover the original AdaGrad and its variant with averaging, and prove a convergence rate of…
Kentaro Matsuura, Junya Honda, Imad El Hanafi, Takashi Sozu + 1 more
'Kentaro Sakamaki'] Estimation of the dose-response curve for efficacy and subsequent selection of an appropriate dose in phase II trials are important processes in drug development. Various methods have been investigated to estimate dose-response curves. Generally, these methods are used with equal allocation of…
Aaron Defazio, Baoyu Zhou, Lin Xiao
The classical AdaGrad method adapts the learning rate by dividing by the square root of a sum of squared gradients. Because this sum on the denominator is increasing, the method can only decrease step sizes over time, and requires a learning rate scaling hyper-parameter to be carefully tuned. To overcome this…
Miaomiao Liu, Dan Yao, Zhigang Liu, Jingfeng Guo + 1 more
An improved Adam optimization algorithm combining adaptive coefficients and composite gradients based on randomized block coordinate descent is proposed to address issues of the Adam algorithm such as slow convergence, the tendency to miss the global optimal solution, and the ineffectiveness of processing…
Reham Elshamy, Osama Abu-Elnasr, Mohamed Elhoseny, Samir Elmougy
Optimizers are the bottleneck of the training process of any Convolutionolution neural networks (CNN) model. One of the critical steps when work on CNN model is choosing the optimal optimizer to solve a specific problem. Recent challenge in nowadays researches is building new versions of traditional CNN optimizers that…
Kosuke Hamazaki, Hiroyoshi Iwata, Koji Tsuda
Differentiable programming frameworks like PyTorch and JAX revolutionized biological modeling. A foremost merit is that multiple components programmed separately can be put together so that the parameters are jointly optimized. Despite its proven value in agricultural applications, existing breeding simulators are…
Yuxing Liu, Rui Pan, Tong Zhang
Adaptive gradient algorithms have been widely adopted in training large-scale deep neural networks, especially large foundation models. Despite their huge success in practice, their theoretical advantages over stochastic gradient descent (SGD) have not been fully understood, especially in the large batch-size setting…
Jianghui Liu, Baozhu Li, Yangfan Zhou, Xuhui Zhao + 2 more
'Mingchuan Zhang'] Adaptive algorithms are widely used because of their fast convergence rate for training deep neural networks (DNNs). However, the training cost becomes prohibitively expensive due to the computation of the full gradient when training complicated DNN. To reduce the computational cost, we present a…
Reham Elshamy, Osama Abu-Elnasr, Mohamed Elhoseny, Samir Elmougy
There are several methods that have been discovered to improve the performance of Deep Learning (DL). Many of these methods reached the best performance of their models by tuning several parameters such as Transfer Learning, Data augmentation, Dropout, and Batch Normalization, while other selects the best optimizer and…
Bastien Batardière, Julien Chiquet, Joon Yeong Kwon
—For finite-sum optimization, variance-reduced gradient methods (VR) compute at each iteration the gradient of a single function (or of a mini-batch), and yet achieve faster convergence than SGD thanks to a carefully crafted lower-variance stochastic gradient estimator that reuses past gradients. Another important line…
Authors not listed
For applications in gas sensing, purification, and capture, we often wish to search a large set of metal-organic frameworks (MOFs) for the top-K in terms of their Henry coefficient of an adsorbate. A molecular simulation to predict the Henry coefficient of a MOF constitutes a Monte Carlo integration where each sample…
Hua-Dong Xiong, Li Ji-An, Robert C. Wilson, Marcelo G. Mattar
A hallmark of intelligence is the ability to adapt behavior to changing environments, which requires adapting one’s own learning strategies. This phenomenon is known as learning to learn or meta-learning. Although well established in humans and animals, a computational framework that characterizes how biological agents…
Kushal Chakrabarti, Mayank Baranwal
Methods under PL Inequality Authors: ['Kushal Chakrabarti' 'Mayank Baranwal'] Abstract. Adaptive gradient-descent optimizers are the standard choice for training neural network models. Despite their faster convergence than gradient-descent and remarkable performance in practice, the adaptive optimizers are not as well…
Finlay Clark, Graeme Robb, Daniel Cole, Julien Michel
Alchemical absolute binding free energy (ABFE) calculations have substantial potential in drug discovery, but are often prohibitively computationally expensive. To unlock their potential, efficient automated ABFE workflows are required to reduce both computational cost and human intervention. We present a…
Huimin Gao, Qingtao Wu, Xuhui Zhao, Junlong Zhu + 2 more
'Petros Daras'] Federated learning is served as a novel distributed training framework that enables multiple clients of the internet of things to collaboratively train a global model while the data remains local. However, the implement of federated learning faces many problems in practice, such as the large number of…
Authors not listed
Finding the most stable adsorption geometry of a flexible molecule on a catalytic surface remains a key challenge due to the high dimensionality and ruggedness of the potential energy surface. We present a Gradient-Enhanced Genetic Algorithm (GE-GA) for the global optimization of adsorbate–surface configurations…
Victor Geadah, Stefan Horoi, Giancarlo Kerg, Guy Wolf + 1 more
Neurons in the brain have rich and adaptive input-output properties. Features such as heterogeneous f-I curves and spike frequency adaptation are known to place single neurons in optimal coding regimes when facing changing stimuli. Yet, it is still unclear how brain circuits exploit single-neuron flexibility, and how…
Charles Micou, Timothy O’Leary
Neural representations of familiar environments and mastered tasks continue to change despite no further refinements to task performance or encoding efficiency. Downstream brain regions that depend on a steady supply of information from a neural population subject to this representational drift face a challenge: they…
Romik Ghosh, Dana Mastrovito, Stefan Mihalas
The human brain readily learns tasks in sequence without forgetting previous ones. Artificial neural networks (ANNs), on the other hand, need to be modified to achieve similar performance. While effective, many algorithms that accomplish this are based on weight importance methods that do not correspond to biological…
John Knight
Latent Factor Analysis via Dynamical Systems (LFADS) is a powerful variational autoencoder for inferring neural population dynamics from spike train data. However, LFADS suffers from pos-terior collapse, where the learned posterior collapses to the prior, eliminating meaningful latent representations. Current solutions…