Abstract
Policy gradient methods for approximate optimal control and reinforcement learning fix parameterized form of the controller and then perform gradient descent on the cost-to-go function. In reinforcement learning for stochastic state-feedback problems, it has been shown that the natural gradient of the cost-to-go function can be approximated via samples of the state and step-cost, using no information about the plant model. There, the natural gradient is the gradient with respect to the Riemannian metric defined by the Fisher information matrix of the controller parameters. We give a general method for approximating the natural gradient for nonlinear output-feedback stochastic control problems with dynamic controllers. For linear systems, we give explicit formulas to compute the natural gradient when plant matrices are known, in both state and output feedback cases.
| Original language | English (US) |
|---|---|
| Title of host publication | 2016 IEEE 55th Conference on Decision and Control, CDC 2016 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 1984-1989 |
| Number of pages | 6 |
| ISBN (Electronic) | 9781509018376 |
| DOIs | |
| State | Published - Dec 27 2016 |
| Event | 55th IEEE Conference on Decision and Control, CDC 2016 - Las Vegas, United States Duration: Dec 12 2016 → Dec 14 2016 |
Publication series
| Name | 2016 IEEE 55th Conference on Decision and Control, CDC 2016 |
|---|
Other
| Other | 55th IEEE Conference on Decision and Control, CDC 2016 |
|---|---|
| Country/Territory | United States |
| City | Las Vegas |
| Period | 12/12/16 → 12/14/16 |
Bibliographical note
Publisher Copyright:© 2016 IEEE.
Fingerprint
Dive into the research topics of 'Natural gradients for state and output feedback control'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS