13 papers · ranked by Valyu relevance
Chang-Ock Lee, Youngkyu Lee, Jong-Ho Park
- Abstract. As deep neural networks (DNNs) become deeper, the training time increases. In this perspective, multi-GPU parallel computing has become a key tool in accelerating the training of DNNs. In this paper, we introduce a novel methodology to construct a parallel neural network that can utilize multiple GPUs…
Milan Curcic
This paper describes neural-fortran, a parallel Fortran framework for neural networks and deep learning. It features a simple interface to construct feed-forward neural networks of arbitrary structure and size, several activation functions, and stochastic gradient descent as the default optimization algorithm.…
Daniela Kalwarowskyj, Erich Schikuta
This paper describes the design and implementation of parallel neural networks (PNNs) with the novel programming language Golang. We follow in our approach the classical Single-Program Multiple-Data (SPMD) model where a PNN is composed of several sequential neural networks, which are trained with a proportional share…
Ludvig Ericson, Rendani Mbuvha
—Artificial Neural Networks (ANNs) have received increasing attention in recent years with applications that span a wide range of disciplines including vital domains such as medicine, network security and autonomous transportation. However, neural network architectures are becoming increasingly complex and with an…
Guang Ping He
We find experimentally that when artificial neural networks are connected in parallel and trained together, they display the following properties. (i) When the parallel-connected neural network (PNN) is optimized, each sub-network in the connection is not optimized. (ii) The contribution of an inferior sub-network to…
Amin Totounferoush, Neda Ebrahimi Pour, Sabine Roller, Miriam Mehl
—In this work1 , we present a parallel scheme for machine learning of partial differential equations. The scheme is based on the decomposition of the training data corresponding to spatial subdomains, where an individual neural network is assigned to each data subset. Message Passing Interface (MPI) is used for…
Yang You, Aydın Buluç, James Demmel
e speed of deep neural networks training has become a big bottleneck of deep learning research and development. For example, training GoogleNet by ImageNet dataset on one Nvidia K20 GPU needs 21 days [11]. To speed up the training process, the current deep learning systems heavily rely on the hardware accelerators.…
Elena Agliari, Andrea Alessandrelli, Adriano Barra, Federico Ricci‐Tersenghi
'Federico Ricci‐Tersenghi'] Abstract: A modern challenge of Artificial Intelligence is learning multiple patterns at once (i.e. parallel learning). While this can not be accomplished by standard Hebbian associative neural networks, in this paper we show how the Multitasking Hebbian Network (a variation on theme of the…
Andrew Simpson
—Although deep neural networks (DNN) are able to scale with direct advances in computational power (e.g., memory and processing speed), they are not well suited to exploit the recent trends for parallel architectures. In particular, gradient descent is a sequential process and the resulting serial dependencies mean…
Salem Alqahtani, Murat Demirbaş
—Deep learning has permeated through many aspects of computing/processing systems in recent years. While distributed training architectures/frameworks are adopted for training large deep learning models quickly, there has not been a systematic study of the communication bottlenecks of these architectures and their…
Antônio H. Ribeiro, Luís A. Aguirre
Neural network models for dynamic systems can be trained either in parallel or in series-parallel configurations. Influenced by early arguments, several papers justify the choice of series-parallel rather than parallel configuration claiming it has a lower computational cost, better stability properties during training…
Mohammad Dehghani, Zahra Yazdanparast
Artificial intelligence has made remarkable progress in handling complex tasks, thanks to advances in hardware acceleration and machine learning algorithms. However, to acquire more accurate outcomes and solve more complex issues, algorithms should be trained with more data. Processing this huge amount of data could be…
Serpen, Gursel
We are proposing fully parallel and maximally distributed hardware realization of a generic neurocomputing system. More specifically, the proposal relates to the wireless sensor networks technology to serve as a massively parallel and fully distributed hardware platform to implement and realize artificial neural…