Past talk · AI/ML Seminar Series
Curiously effective ensemble and double-oracle reinforcement-learning methods
Assistant Professor, Department of Computer Science, University of California, Irvine
- Date & time
- Monday, January 10, 2022 · 1:00 PM
- Location
- Online (live stream)
Abstract
Ensemble methods for reinforcement learning represent model uncertainty and use it to guide exploration and reduce value estimation bias. We present MeanQ, a very simple ensemble method with improved performance that reduces estimation variance enough to operate without a stabilizing target network — curiously, it is theoretically almost equivalent to a non-ensemble method it significantly outperforms. In adversarial environments, double-oracle (DO) methods grow a population of policies by iteratively adding best responses. We present XDO, a DO algorithm that exploits the game's sequential structure to exponentially reduce the worst-case population size.
About the speaker
Roy Fox is an Assistant Professor and director of the Intelligent Dynamics Lab in the Department of Computer Science at UCI. He was previously a postdoc in UC Berkeley's BAIR, RISELab, and AUTOLAB. His research interests include theory and applications of reinforcement learning, algorithmic game theory, information theory, and robotics, with a current focus on structure, exploration, and optimization in deep RL and imitation learning.