BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//UC Irvine//CML Seminars//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Curiously effective ensemble and double-oracle reinforcement-l
 earning methods
X-WR-TIMEZONE:America/Los_Angeles
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:20070311T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:20071104T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
UID:2022-01-10-roy-fox@cml.ics.uci.edu
DTSTAMP:20220110T000000Z
SEQUENCE:57690
DTSTART;TZID=America/Los_Angeles:20220110T130000
DTEND;TZID=America/Los_Angeles:20220110T140000
SUMMARY:[CML Seminar] Roy Fox: Curiously effective ensemble and double-orac
 le reinforcement-learning methods
LOCATION:Online (live stream)
DESCRIPTION:Roy Fox\, Assistant Professor\, Department of Computer Science\
 , University of California\, Irvine\n\nTitle: Curiously effective ensemble
  and double-oracle reinforcement-learning methods\n\nAbstract: Ensemble me
 thods for reinforcement learning represent model uncertainty and use it to
  guide exploration and reduce value estimation bias. We present MeanQ\, a 
 very simple ensemble method with improved performance that reduces estimat
 ion variance enough to operate without a stabilizing target network — cu
 riously\, it is theoretically almost equivalent to a non-ensemble method i
 t significantly outperforms. In adversarial environments\, double-oracle (
 DO) methods grow a population of policies by iteratively adding best respo
 nses. We present XDO\, a DO algorithm that exploits the game's sequential 
 structure to exponentially reduce the worst-case population size.\n\nhttps
 ://cml.ics.uci.edu/seminars/2022-01-10-roy-fox
X-ALT-DESC;FMTTYPE=text/html:<html><body><b>Roy Fox</b>\, Assistant Profess
 or\, Department of Computer Science\, University of California\, Irvine<br
 ><br><b>Title:</b> Curiously effective ensemble and double-oracle reinforc
 ement-learning methods<br><br><b>Abstract:</b> Ensemble methods for reinfo
 rcement learning represent model uncertainty and use it to guide explorati
 on and reduce value estimation bias. We present MeanQ\, a very simple ense
 mble method with improved performance that reduces estimation variance eno
 ugh to operate without a stabilizing target network — curiously\, it is 
 theoretically almost equivalent to a non-ensemble method it significantly 
 outperforms. In adversarial environments\, double-oracle (DO) methods grow
  a population of policies by iteratively adding best responses. We present
  XDO\, a DO algorithm that exploits the game's sequential structure to exp
 onentially reduce the worst-case population size.<br><br><a href="https://
 cml.ics.uci.edu/seminars/2022-01-10-roy-fox">https://cml.ics.uci.edu/semin
 ars/2022-01-10-roy-fox</a></body></html>
URL:https://cml.ics.uci.edu/seminars/2022-01-10-roy-fox
END:VEVENT
END:VCALENDAR
