DOE OSTI · 2325260
Exploration with Scalable Gaussian Process Reinforcement Learning
Abstract
Exploration is a challenging problem in reinforcement learning (RL), especially in environments with sparse rewards. Quantifying and utilizing the parametric uncertainty has been shown to be paramount for successful exploration [Osband et al., 2018]. Bayesian, or approximately Bayesian, methods present a principled means of estimating the parametric uncertainty in RL problems. Gaussian processes, nonparametric Bayesian models, are often impractical due to poor scalability and computational bottlenecks. We introduce a scalable Gaussian process RL (GPRL) method which directly induces sparsity in the covariance matrix to facilitate faster computation. This is a departure from previous GPRL methods which instead rely on data reduction and subsampling. We compare various covariance-based exploration techniques (Thompson sampling, upper confidence bound, and probabilistic maximum variance) which leverage our scalable GP framework in sparse reward environments. Finally, we show favorable comparison against the bootstrapped deep Q-Network.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Miller, Caleb J., Soper, Braden C., Muyskens, Amanda, Priest, Benjamin W., Schneider, Michael D., Merl, Dan M.. 2024-03-07. Exploration with Scalable Gaussian Process Reinforcement Learning. https://doi.org/10.2172/2325260
Cite the original work for its findings. Save a collection to share your selection of sources.