Results 31 to 40 of about 22,588 (154)
Understanding and Optimizing Asynchronous Low-Precision Stochastic Gradient Descent [PDF]
Chris Re, Kunle Olukotun
exaly +2 more sources
The effective noise of stochastic gradient descent
Abstract Stochastic gradient descent (SGD) is the workhorse algorithm of deep learning technology. At each step of the training phase, a mini batch of samples is drawn from the training dataset and the weights of the neural network are adjusted according to the performance on this specific subset of examples.
Mignacco, Francesca +1 more
openaire +2 more sources
Pipelined Stochastic Gradient Descent with Taylor Expansion
Stochastic gradient descent (SGD) is an optimization method typically used in deep learning to train deep neural network (DNN) models. In recent studies for DNN training, pipeline parallelism, a type of model parallelism, is proposed to accelerate SGD ...
Bongwon Jang, Inchul Yoo, Dongsuk Yook
doaj +1 more source
Stochastic Gradient Descent in Continuous Time [PDF]
Stochastic gradient descent in continuous time (SGDCT) provides a computationally efficient method for the statistical learning of continuous-time models, which are widely used in science, engineering, and finance. The SGDCT algorithm follows a (noisy) descent direction along a continuous stream of data.
Justin A. Sirignano +1 more
openaire +2 more sources
Byzantine Stochastic Gradient Descent
This paper studies the problem of distributed stochastic optimization in an adversarial setting where, out of the $m$ machines which allegedly compute stochastic gradients every iteration, an $α$-fraction are Byzantine, and can behave arbitrarily and adversarially.
Alistarh, Dan-Adrian +2 more
openaire +4 more sources
Scaling transition from momentum stochastic gradient descent to plain stochastic gradient descent
The plain stochastic gradient descent and momentum stochastic gradient descent have extremely wide applications in deep learning due to their simple settings and low computational complexity. The momentum stochastic gradient descent uses the accumulated gradient as the updated direction of the current parameters, which has a faster training speed ...
Kun Zeng +3 more
openaire +2 more sources
Natural Evolutionary Gradient Descent Strategy for Variational Quantum Algorithms
Recent research has demonstrated that parametric quantum circuits (PQCs) are affected by gradients that progressively vanish to zero as a function of the number of qubits.
Jianshe Xie +4 more
doaj +1 more source
Randomized Stochastic Gradient Descent Ascent
An increasing number of machine learning problems, such as robust or adversarial variants of existing algorithms, require minimizing a loss function that is itself defined as a maximum. Carrying a loop of stochastic gradient ascent (SGA) steps on the (inner) maximization problem, followed by an SGD step on the (outer) minimization, is known as Epoch ...
Othmane Sebbouh +2 more
openaire +3 more sources
Stochastic gradient-free descents
In this paper we propose stochastic gradient-free methods and accelerated methods with momentum for solving stochastic optimization problems. All these methods rely on stochastic directions rather than stochastic gradients. We analyze the convergence behavior of these methods under the mean-variance framework, and also provide a theoretical analysis ...
Xiaopeng Luo, Xin Xu 0006
openaire +2 more sources
On the discrepancy principle for stochastic gradient descent
Abstract Stochastic gradient descent (SGD) is a promising numerical method for solving large-scale inverse problems. However, its theoretical properties remain largely underexplored in the lens of classical regularization theory. In this note, we study the classical discrepancy principle, one of the most popular a posteriori choice rules,
Tim Jahn, Bangti Jin
openaire +4 more sources

