14 papers · ranked by Valyu relevance
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer + 17 more
'Gregory Chanan' 'Trevor Killeen' 'Zeming Lin' 'Natalia Gimelshein' 'Luca Antiga' 'Alban Desmaison' 'Andreas Köpf' 'Edward Z. Yang' 'Zach DeVito' 'Martin Raison' 'Alykhan Tejani' 'Sasank Chilamkurthy' 'Benoit Steiner' 'Lu Fang' 'Junjie Bai' 'Soumith Chintala'] Deep learning frameworks have often focused on either…
Yueming Hao, Xu Zhao, Bin Bao, David Berard + 3 more
'Adnan Aziz' 'Xu Liu'] Deep learning (DL) has been a revolutionary technique in various domains. To facilitate the model development and deployment, many deep learning frameworks are proposed, among which Py-Torch is one of the most popular solutions. The performance of ecosystem around PyTorch is critically important…
Abhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava Basu
CUDA Graphs — a recent hardware feature introduced for NVIDIA GPUs — aim to reduce CPU launch overhead by capturing and launching a series of GPU tasks (kernels) as a DAG. However, deploying CUDA Graphs faces several challenges today due to the static structure of a graph. It also incurs performance overhead due to…
Zakariya Ba Alawi
—This paper presents a comprehensive comparative survey of TensorFlow and PyTorch, the two leading deep learning frameworks, focusing on their usability, performance, and deployment trade-offs. We review each framework's programming paradigm and developer experience, contrasting TensorFlow's graph-based (now optionally…
Seung Won Min, Wu Kun, Sitao Huang, Mert Hidayetoğlu + 4 more
'Eiman Ebrahimi' 'Deming Chen' 'Wen‐mei Hwu'] With the increasing adoption of graph neural networks (GNNs) in the machine learning community, GPUs have become an essential tool to accelerate GNN training. However, training GNNs on very large graphs that do not fit in GPU memory is still a challenging task. Unlike…
Tomasz Kornuta
Access to vast amounts of data along with affordable computational power stimulated the reincarnation of neural networks. The progress could not be achieved without adequate software tools, lowering the entry bar for the next generations of researchers and developers. The paper introduces PyTorchPipe (PTP), a framework…
Jonas Rauber, Matthias Bethge, Wieland Brendel
EagerPy is a Python framework that lets you write code that automatically works natively with PyTorch, TensorFlow, JAX, and NumPy. Library developers no longer need to choose between supporting just one of these frameworks or reimplementing the library for each framework and dealing with code duplication. Users of such…
Ho Young Jhoo, Sehoon Kim, Woosung Song, Kyuyeon Park + 2 more
'Dong-Kwon Lee' 'Kwangkeun Yi'] We present an automatic static analyzer PyTea that detects tensorshape errors in PyTorch code. The tensor-shape error is critical in the deep neural net code; much of the training cost and intermediate results are to be lost once a tensor shape mismatch occurs in the midst of the…
Bojan Nikolic
—I show that a software framework intended primarily for training of neural networks, PyTorch, is easily applied to a general function minimisation problem in science. The qualities of PyTorch of ease-of-use and very high efficiency are found to be applicable in this domain and lead to two orders of magnitude…
Manu Joseph
— In spite of showing unreasonable effectiveness in modalities like Text and Image, Deep Learning has always lagged Gradient Boosting in tabular data—both in popularity and performance. But recently there have been newer models created specifically for tabular data, which is pushing the performance bar. But popularity…
Zachary DeVito, Jason Ansel, Will Constable, Michael Suo + 2 more
'Ailing Zhang' 'Kim Hazelwood'] Python has become the de-facto language for training deep neural networks, coupling a large suite of scientific computing libraries with efficient libraries for tensor computation such as PyTorch (Paszke et al., 2019) or TensorFlow (Abadi et al., 2016). However, when models are used for…
Charlene Yang, Yunsong Wang, Thorsten Kurth, Steven Farrell + 1 more
'Samuel Williams'] Abstract—This paper presents a practical methodology for collecting performance data necessary to conduct hierarchical Roofline analysis on NVIDIA GPUs. It discusses the extension of the Empirical Roofline Toolkit for broader support of a range of data precisions and Tensor Core support and…
Thomas Bartz–Beielstein
The goal of hyperparameter tuning (or hyperparameter optimization) is to optimize the hyperparameters to improve the performance of the machine or deep learning model. spotPython ("Sequential Parameter Optimization Toolbox in Python") is the Python version of the well-known hyperparameter tuner SPOT, which has been…
Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, James Bergstra + 5 more
'Ian J. Goodfellow' 'Arnaud Bergeron' 'Nicolas Bouchard' 'David Warde-Farley' 'Yoshua Bengio'] Theano is a linear algebra compiler that optimizes a user's symbolically-specified mathematical computations to produce efficient low-level implementations. In this paper, we present new features and efficiency improvements…