BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//UC Irvine//CML Seminars//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:AppWorld: Reliable Evaluation of Interactive Agents in a World
  of Apps and People
X-WR-TIMEZONE:America/Los_Angeles
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:20070311T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:20071104T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
UID:2024-10-07-harsh-trivedi@cml.ics.uci.edu
DTSTAMP:20241007T000000Z
SEQUENCE:57690
DTSTART;TZID=America/Los_Angeles:20241007T130000
DTEND;TZID=America/Los_Angeles:20241007T140000
SUMMARY:[CML Seminar] Harsh Trivedi: AppWorld: Reliable Evaluation of Inter
 active Agents in a World of Apps and People
LOCATION:Donald Bren Hall 4011
DESCRIPTION:Harsh Trivedi\, PhD Student\, Department of Computer Science\, 
 Stony Brook University\n\nTitle: AppWorld: Reliable Evaluation of Interact
 ive Agents in a World of Apps and People\n\nAbstract: We envision a world 
 where AI agents (assistants) are widely used for complex tasks in our digi
 tal and physical worlds and are broadly integrated into our society. To mo
 ve towards such a future\, we need an environment for a robust evaluation 
 of agents' capability\, reliability\, and trustworthiness. In this talk\, 
 I'll introduce AppWorld\, which is a step towards this goal in the context
  of day-to-day digital tasks. AppWorld is a high-fidelity simulated world 
 of people and their digital activities on nine apps like Amazon\, Gmail\, 
 and Venmo. On top of this fully controllable world\, we build a benchmark 
 of complex day-to-day tasks such as splitting Venmo bills with roommates\,
  which agents have to solve via interactive coding and API calls. Our benc
 hmarking evaluations show that even the best LLMs\, like GPT-4o\, can only
  solve ~30% of such tasks\, highlighting the challenging nature of the App
 World benchmark.\n\nhttps://cml.ics.uci.edu/seminars/2024-10-07-harsh-triv
 edi
X-ALT-DESC;FMTTYPE=text/html:<html><body><b>Harsh Trivedi</b>\, PhD Student
 \, Department of Computer Science\, Stony Brook University<br><br><b>Title
 :</b> AppWorld: Reliable Evaluation of Interactive Agents in a World of Ap
 ps and People<br><br><b>Abstract:</b> We envision a world where AI agents 
 (assistants) are widely used for complex tasks in our digital and physical
  worlds and are broadly integrated into our society. To move towards such 
 a future\, we need an environment for a robust evaluation of agents' capab
 ility\, reliability\, and trustworthiness. In this talk\, I'll introduce A
 ppWorld\, which is a step towards this goal in the context of day-to-day d
 igital tasks. AppWorld is a high-fidelity simulated world of people and th
 eir digital activities on nine apps like Amazon\, Gmail\, and Venmo. On to
 p of this fully controllable world\, we build a benchmark of complex day-t
 o-day tasks such as splitting Venmo bills with roommates\, which agents ha
 ve to solve via interactive coding and API calls. Our benchmarking evaluat
 ions show that even the best LLMs\, like GPT-4o\, can only solve ~30% of s
 uch tasks\, highlighting the challenging nature of the AppWorld benchmark.
 <br><br><a href="https://cml.ics.uci.edu/seminars/2024-10-07-harsh-trivedi
 ">https://cml.ics.uci.edu/seminars/2024-10-07-harsh-trivedi</a></body></ht
 ml>
URL:https://cml.ics.uci.edu/seminars/2024-10-07-harsh-trivedi
END:VEVENT
END:VCALENDAR
