18 papers · ranked by Valyu relevance
Dominic Masters, Carlo Luschi
Modern deep neural network training is typically based on mini-batch stochastic gradient optimization. While the use of large mini-batches increases the available computational parallelism, small batch training has been shown to provide improved generalization performance and allows a significantly smaller memory…
Paul Stapor, Leonard Schmiester, Christoph Wierling, Simon Merkt + 4 more
Quantitative dynamic models are widely used to study cellular signal processing. A critical step in modelling is the estimation of unknown model parameters from experimental data. As model sizes and datasets are steadily growing, established parameter optimization approaches for mechanistic models become…
Paul Stapor, Leonard Schmiester, Christoph Wierling, Bodo M.H. Lange + 2 more
Quantitative dynamical models are widely used to study cellular signal processing. A critical step in modeling is the estimation of unknown model parameters from experimental data. As model sizes and datasets are steadily growing, established parameter optimization approaches for mechanistic models become…
Sukey Nakasima-López, Juan R. Castro, Mauricio A. Sanchez, Olivia Mendoza + 2 more
'Olivia Mendoza' 'Antonio Rodríguez-Díaz' 'Jie Zhang'] Due to the rapid technological evolution and communications accessibility, data generated from different sources of information show an exponential growth behavior. That is, volume of data samples that need to be analyzed are getting larger, so the methods for its…
Scott Sievert
Mini-batch stochastic gradient descent (SGD) and variants thereof approximate the objective function's gradient with a small number of training examples, aka the batch size. Small batch sizes require little computation for each model update but can yield high-variance gradient estimates, which poses some challenges for…
Dong Yin, Ashwin Pananjady, Max W. Y. Lam, Dimitris Papailiopoulos + 2 more
'Kannan Ramchandran' 'Peter L. Bartlett'] It has been experimentally observed that distributed implementations of mini-batch stochastic gradient descent (SGD) algorithms exhibit speedup saturation and decaying generalization ability beyond a particular batch-size. In this work, we present an analysis hinting that high…
Subin Sahayam, John Zakkam, Umarani Jayaraman
In deep learning, mini-batch training is commonly used to optimize network parameters. However, the traditional mini-batch method may not learn the under-represented samples and complex patterns in the data, leading to a longer time for generalization. To address this problem, a variant of the traditional algorithm has…
Xue Wang, Yinghan Chen, Shiyu Wang
Cognitive diagnostic models (CDMs) provide fine-grained diagnostic feedback by modeling the relationship between latent attributes and item responses. Two key components required for CDM implementation are the Q-matrix, which links items to attributes, and the attribute hierarchy, which defines prerequisite…
XinYu Piao, DoangJoo Synn, Jooyoung Park, Jong‐Kook Kim
Recent deep learning models are difficult to train using a large batch size, because commodity machines may not have enough memory to accommodate both the model and a large data batch size. The batch size is one of the hyper-parameters used in the training model, and it is dependent on and is limited by the target…
Tom Schaul, Yann LeCun
Recent work has established an empirically successful framework for adapting learning rates for stochastic gradient descent (SGD). This effectively removes all needs for tuning, while automatically reducing learning rates over time on stationary problems, and permitting learning rates to grow appropriately in…
Stephanie C. Hicks, Ruoxi Liu, Yuwei Ni, Elizabeth Purdom + 1 more
Single-cell RNA-Sequencing (scRNA-seq) is the most widely used high-throughput technology to measure genome-wide gene expression at the single-cell level. One of the most common analyses of scRNA-seq data detects distinct subpopulations of cells through the use of unsupervised clustering algorithms. However, recent…
Hyeonseong Choi, Byung Hyun Lee, Se Young Chun, Jaehwan Lee + 1 more
'Elena Loli Piccolomini'] Modern deep neural networks cannot be often trained on a single GPU due to large model size and large data size. Model parallelism splits a model for multiple GPUs, but making it scalable and seamless is challenging due to different information sharing among GPUs with communication overhead.…
Chao Gao, Joshua D. Welch
Recent experimental advances have enabled high-throughput single-cell measurement of gene expression, chromatin accessibility and DNA methylation. We previously used integrative non-negative matrix factorization (iNMF) to jointly learn interpretable low-dimensional representations from multiple single-cell datasets…
Ya Chen, Thomas Seidel, Roxane Axel Jacob, Steffen Hirte + 6 more
The ability to determine and predict metabolically labile atom positions in a molecule (also called “sites of metabolism” or “SoMs”) is of high interest to the design and optimization of bioactive compounds such as drugs, agrochemicals, and cosmetics. In recent years, several in silico models for SoM prediction have…
Michael Bailey, Saeed Moayedpour, Ruijiang Li, Alejandro Corrochano-Navarro + 10 more
A key challenge in drug discovery is to optimize, in silico, various absorption and affinity properties of small molecules. One strategy that was proposed for such optimization process is active learning. In active learning molecules are selected for testing based on their likelihood of improving model performance. To…
Mohamed Soudy, Yasmine Afify, Nagwa Badr, Yilun Shang
Image understanding and scene classification are keystone tasks in computer vision. The development of technologies and profusion of existing datasets open a wide room for improvement in the image classification and recognition research area. Notwithstanding the optimal performance of exiting machine learning models in…
Authors not listed
Large Language Models (LLMs) based on transformer architectures excel at internet-scale tasks. However, real-world scientific scenarios—such as synthetic chemistry laboratories and autonomous experimental setups—typically involve incremental data generation in batches as new chemical reactions are conducted, unlike…
Simon Viet Johansson, Hampus Gummesson Svensson, Esben Bjerrum, Alexander Schliep + 3 more
Computer aided synthesis planning is a rapidly growing field for suggesting synthetic routes for molecules of interest. The methods used are usually dependent on access to large datasets for training, but with a finite experimental budget there are limitations on how much data can be obtained from experiments. Active…