Results 21 to 30 of about 22,826 (253)

Pipelined Stochastic Gradient Descent with Taylor Expansion

open access: yesApplied Sciences, 2023
Stochastic gradient descent (SGD) is an optimization method typically used in deep learning to train deep neural network (DNN) models. In recent studies for DNN training, pipeline parallelism, a type of model parallelism, is proposed to accelerate SGD ...
Bongwon Jang, Inchul Yoo, Dongsuk Yook
doaj   +1 more source

Byzantine Stochastic Gradient Descent

open access: yesCoRR, 2018
This paper studies the problem of distributed stochastic optimization in an adversarial setting where, out of the $m$ machines which allegedly compute stochastic gradients every iteration, an $α$-fraction are Byzantine, and can behave arbitrarily and adversarially.
Alistarh, Dan-Adrian   +2 more
openaire   +4 more sources

Scaling transition from momentum stochastic gradient descent to plain stochastic gradient descent

open access: yesCoRR, 2021
The plain stochastic gradient descent and momentum stochastic gradient descent have extremely wide applications in deep learning due to their simple settings and low computational complexity. The momentum stochastic gradient descent uses the accumulated gradient as the updated direction of the current parameters, which has a faster training speed ...
Kun Zeng   +3 more
openaire   +3 more sources

Natural Evolutionary Gradient Descent Strategy for Variational Quantum Algorithms

open access: yesIntelligent Computing, 2023
Recent research has demonstrated that parametric quantum circuits (PQCs) are affected by gradients that progressively vanish to zero as a function of the number of qubits.
Jianshe Xie   +4 more
doaj   +1 more source

Stochastic gradient-free descents

open access: yesCoRR, 2019
In this paper we propose stochastic gradient-free methods and accelerated methods with momentum for solving stochastic optimization problems. All these methods rely on stochastic directions rather than stochastic gradients. We analyze the convergence behavior of these methods under the mean-variance framework, and also provide a theoretical analysis ...
Xiaopeng Luo, Xin Xu 0006
openaire   +3 more sources

On the discrepancy principle for stochastic gradient descent

open access: yesInverse Problems, 2020
Abstract Stochastic gradient descent (SGD) is a promising numerical method for solving large-scale inverse problems. However, its theoretical properties remain largely underexplored in the lens of classical regularization theory. In this note, we study the classical discrepancy principle, one of the most popular a posteriori choice rules,
Tim Jahn, Bangti Jin
openaire   +5 more sources

Benign Underfitting of Stochastic Gradient Descent

open access: yesAdvances in Neural Information Processing Systems 35, 2022
We study to what extent may stochastic gradient descent (SGD) be understood as a "conventional" learning rule that achieves generalization performance by obtaining a good fit to training data. We consider the fundamental stochastic convex optimization framework, where (one pass, without-replacement) SGD is classically known to minimize the population ...
Tomer Koren   +3 more
openaire   +3 more sources

Preconditioned Stochastic Gradient Descent [PDF]

open access: yesIEEE Transactions on Neural Networks and Learning Systems, 2018
13 pages, 9 figures. To appear in IEEE Transactions on Neural Networks and Learning Systems.
openaire   +3 more sources

Featured Hybrid Recommendation System Using Stochastic Gradient Descent

open access: yesInternational Journal of Networked and Distributed Computing (IJNDC), 2021
Beside cold-start and sparsity, developing incremental algorithms emerge as interesting research to recommendation system in real-data environment. While hybrid system research is insufficient due to the complexity in combining various source of each ...
Si Thin Nguyen   +3 more
doaj   +1 more source

Randomized Stochastic Gradient Descent Ascent

open access: yesCoRR, 2021
An increasing number of machine learning problems, such as robust or adversarial variants of existing algorithms, require minimizing a loss function that is itself defined as a maximum. Carrying a loop of stochastic gradient ascent (SGA) steps on the (inner) maximization problem, followed by an SGD step on the (outer) minimization, is known as Epoch ...
Othmane Sebbouh   +2 more
openaire   +3 more sources

Home - About - Disclaimer - Privacy