Skip to main content
← Back to seminars

Past talk · AI/ML Seminar Series

Curiously effective ensemble and double-oracle reinforcement-learning methods

Roy Fox

Assistant Professor, Department of Computer Science, University of California, Irvine

Date & time
Monday, January 10, 2022 · 1:00 PM
Location
Online (live stream)

Abstract

Ensemble methods for reinforcement learning represent model uncertainty and use it to guide exploration and reduce value estimation bias. We present MeanQ, a very simple ensemble method with improved performance that reduces estimation variance enough to operate without a stabilizing target network — curiously, it is theoretically almost equivalent to a non-ensemble method it significantly outperforms. In adversarial environments, double-oracle (DO) methods grow a population of policies by iteratively adding best responses. We present XDO, a DO algorithm that exploits the game's sequential structure to exponentially reduce the worst-case population size.

About the speaker

Roy Fox is an Assistant Professor and director of the Intelligent Dynamics Lab in the Department of Computer Science at UCI. He was previously a postdoc in UC Berkeley's BAIR, RISELab, and AUTOLAB. His research interests include theory and applications of reinforcement learning, algorithmic game theory, information theory, and robotics, with a current focus on structure, exploration, and optimization in deep RL and imitation learning.