Search · four archives
Search · four archives
23 papers · ranked by Valyu relevance
Yunji Yang, Yonggi Hong, Jaehyun Park, Leopoldo Angrisani
In this paper, efficient gradient updating strategies are developed for the federated learning when distributed clients are connected to the server via a wireless backhaul link. Specifically, a common convolutional neural network (CNN) module is shared for all the distributed clients and it is trained through the…
Amirhessam Tahmassebi, Amir H. Gandomi, Simon Fong, Anke Meyer-Baese + 2 more
'Simon Y. Foo' 'Ivan Olier'] In this study, a multi-stage optimization procedure is proposed to develop deep neural network models which results in a powerful deep learning pipeline called intelligent deep learning (iDeepLe). The proposed pipeline is then evaluated by a challenging real-world problem, the modeling of…
Jiawei Zhang
In this paper, we aim at providing an introduction to the gradient descent based optimization algorithms for learning deep neural network models. Deep learning models involving multiple nonlinear projection layers are very challenging to train. Nowadays, most of the deep learning model training still relies on the back…
Yash Ganpat Sawant
Adaptive rank allocation for LoRA, allocating more parameters to important layers and fewer to unimportant ones, consistently improves efficiency under supervised fine-tuning (SFT). We investigate whether this success transfers to reinforcement learning, specifically Group Relative Policy Optimization (GRPO). Using…
Cyrille W. Combettes, Christoph Spiegel, Sebastian Pokutta
The complexity in large-scale optimization can lie in both handling the objective function and handling the constraint set. In this respect, stochastic Frank-Wolfe algorithms occupy a unique position as they alleviate both computational burdens, by querying only approximate first-order information from the objective…
Jinghui Chen, Dongruo Zhou, Yiqi Tang, Ziyan Yang + 2 more
'Quanquan Gu'] Adaptive gradient methods, which adopt historical gradient information to automatically adjust the learning rate, despite the nice property of fast convergence, have been observed to generalize worse than stochastic gradient descent (SGD) with momentum in training deep neural networks. This leaves how to…
Miaomiao Liu, Dan Yao, Zhigang Liu, Jingfeng Guo + 1 more
An improved Adam optimization algorithm combining adaptive coefficients and composite gradients based on randomized block coordinate descent is proposed to address issues of the Adam algorithm such as slow convergence, the tendency to miss the global optimal solution, and the ineffectiveness of processing…
Tianyi Chen, Qing Ling, Georgios B. Giannakis
—Network resource allocation shows revived popularity in the era of data deluge and information explosion. Existing stochastic optimization approaches fall short in attaining a desirable cost-delay tradeoff. Recognizing the central role of Lagrange multipliers in network resource allocation, a novel learn-andadapt…
Reham Elshamy, Osama Abu-Elnasr, Mohamed Elhoseny, Samir Elmougy
Optimizers are the bottleneck of the training process of any Convolutionolution neural networks (CNN) model. One of the critical steps when work on CNN model is choosing the optimal optimizer to solve a specific problem. Recent challenge in nowadays researches is building new versions of traditional CNN optimizers that…
Kosuke Hamazaki, Hiroyoshi Iwata, Koji Tsuda
Differentiable programming frameworks like PyTorch and JAX revolutionized biological modeling. A foremost merit is that multiple components programmed separately can be put together so that the parameters are jointly optimized. Despite its proven value in agricultural applications, existing breeding simulators are…
A. R. Flores, Rodrigo C. de Lamare
Rate splitting (RS) systems can better deal with imperfect channel state information at the transmitter (CSIT) than conventional approaches. However, this requires an appropriate power allocation that often has a high computational complexity, which might be inadequate for practical and large systems. To this end…
Kentaro Matsuura, Junya Honda, Imad El Hanafi, Takashi Sozu + 1 more
'Kentaro Sakamaki'] Estimation of the dose-response curve for efficacy and subsequent selection of an appropriate dose in phase II trials are important processes in drug development. Various methods have been investigated to estimate dose-response curves. Generally, these methods are used with equal allocation of…
Reham Elshamy, Osama Abu-Elnasr, Mohamed Elhoseny, Samir Elmougy
There are several methods that have been discovered to improve the performance of Deep Learning (DL). Many of these methods reached the best performance of their models by tuning several parameters such as Transfer Learning, Data augmentation, Dropout, and Batch Normalization, while other selects the best optimizer and…
Kin Gutierrez, Jin Li, Cristian Challú, Artur Dubrawski
Adaptive moment methods have been remarkably successful in deep learning optimization, particularly in the presence of noisy and/or sparse gradients. We further the advantages of adaptive moment techniques by proposing a family of double adaptive stochastic gradient methods DASGrad. They leverage the complementary…
Authors not listed
For applications in gas sensing, purification, and capture, we often wish to search a large set of metal-organic frameworks (MOFs) for the top-K in terms of their Henry coefficient of an adsorbate. A molecular simulation to predict the Henry coefficient of a MOF constitutes a Monte Carlo integration where each sample…
Hua-Dong Xiong, Li Ji-An, Robert C. Wilson, Marcelo G. Mattar
A hallmark of intelligence is the ability to adapt behavior to changing environments, which requires adapting one’s own learning strategies. This phenomenon is known as learning to learn or meta-learning. Although well established in humans and animals, a computational framework that characterizes how biological agents…
Kazunori D Yamada
In the deep learning era, a gradient descent method is the most common method to optimize parameters of neural networks. Among various mathematical optimization methods, a gradient descent method is the most naive method. Although controlling a learning rate of the method is necessary for quick convergence, the…
Finlay Clark, Graeme Robb, Daniel Cole, Julien Michel
Alchemical absolute binding free energy (ABFE) calculations have substantial potential in drug discovery, but are often prohibitively computationally expensive. To unlock their potential, efficient automated ABFE workflows are required to reduce both computational cost and human intervention. We present a…
Huimin Gao, Qingtao Wu, Xuhui Zhao, Junlong Zhu + 2 more
'Petros Daras'] Federated learning is served as a novel distributed training framework that enables multiple clients of the internet of things to collaboratively train a global model while the data remains local. However, the implement of federated learning faces many problems in practice, such as the large number of…
Authors not listed
Finding the most stable adsorption geometry of a flexible molecule on a catalytic surface remains a key challenge due to the high dimensionality and ruggedness of the potential energy surface. We present a Gradient-Enhanced Genetic Algorithm (GE-GA) for the global optimization of adsorbate–surface configurations…
Victor Geadah, Stefan Horoi, Giancarlo Kerg, Guy Wolf + 1 more
Neurons in the brain have rich and adaptive input-output properties. Features such as heterogeneous f-I curves and spike frequency adaptation are known to place single neurons in optimal coding regimes when facing changing stimuli. Yet, it is still unclear how brain circuits exploit single-neuron flexibility, and how…
Vladimír Kunc, Jiří Kléma
Gene expression profiling was made cheaper by the NIH LINCS program that profiles only ~1, 000 selected landmark genes and uses them to reconstruct the whole profile. The D–GEX method employs neural networks to infer the whole profile. However, the original D–GEX can be further significantly improved. We have analyzed…
Authors not listed
Quantum mechanics/molecular mechanics (QM/MM) simulations are crucial for understanding enzymatic reactions, but their accuracy depends heavily on the quantum-mechanical method used. Semiempirical methods offer computational efficiency but often struggle with accuracy in complex systems. This work presents a novel…