Linear Convergence of Entropy-Regularized Natural Policy Gradient with Linear Function Approximation
METADATA ONLY
Loading...
Author / Creator
Date
2022-02-17
Publication Type
Working Paper
ETH Bibliography
yes
Citations
Altmetric
METADATA ONLY
Data
Rights / License
Abstract
Natural policy gradient (NPG) methods with entropy regularization achieve impressive empirical success in reinforcement learning problems with large state-action spaces. However, their convergence properties and the impact of entropy regularization remain elusive in the function approximation regime. In this paper, we establish finite-time convergence analyses of entropy-regularized NPG with linear function approximation under softmax parameterization. In particular, we prove that entropy-regularized NPG with averaging satisfies the \emph{persistence of excitation} condition, and achieves a fast convergence rate of $\tilde{O}(1/T)$ up to a function approximation error in regularized Markov decision processes. This convergence result does not require any a priori assumptions on the policies. Furthermore, under mild regularity conditions on the concentrability coefficient and basis vectors, we prove that entropy-regularized NPG exhibits \emph{linear convergence} up to a function approximation error.
Permanent link
Publication status
published
Editor
Book title
Journal / series
Volume
Pages / Article No.
2106.04096v3
Publisher
Cornell University
Event
Edition / version
3
Methods
Geographic location
Date collected
Date created
Subject
Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML); FOS: Computer and information sciences; FOS: Mathematics
Organisational unit
09729 - He, Niao / He, Niao
Notes
Funding Info about funding
Related publications and datasets
Is new version of: