BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//UC Irvine//CML Seminars//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:What Generative Visual Models Understand (and Don't) About the
  Physical World
X-WR-TIMEZONE:America/Los_Angeles
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:20070311T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:20071104T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
UID:2025-03-17-anand-bhattad@cml.ics.uci.edu
DTSTAMP:20250317T000000Z
SEQUENCE:57690
DTSTART;TZID=America/Los_Angeles:20250317T110000
DTEND;TZID=America/Los_Angeles:20250317T120000
SUMMARY:[CML Seminar] Anand Bhattad: What Generative Visual Models Understa
 nd (and Don't) About the Physical World
LOCATION:Donald Bren Hall 4011
DESCRIPTION:Anand Bhattad\, Research Assistant Professor\, Toyota Technolog
 ical Institute at Chicago\n\nTitle: What Generative Visual Models Understa
 nd (and Don't) About the Physical World\n\nAbstract: Generative visual mod
 els like Stable Diffusion and Sora generate photorealistic images and vide
 os that are nearly indistinguishable from real ones to a naive observer. H
 owever\, their grasp of the physical world remains an open question: Do th
 ey understand 3D geometry\, light\, and object interactions\, or are they 
 mere pixel parrots of their training data? Through systematic probing\, I 
 will demonstrate that these models surprisingly learn fundamental scene pr
 operties — intrinsic images such as surface normals\, depth\, albedo\, a
 nd shading — without explicit supervision\, which enables applications l
 ike image relighting. But I will also show that this knowledge is insuffic
 ient. Careful analysis reveals unexpected failures: inconsistent shadows\,
  multiple vanishing points\, and scenes that defy basic physics. All these
  findings suggest these models excel at local texture synthesis but strugg
 le with global reasoning: a crucial gap between imitation and true underst
 anding. I will then conclude by outlining a path toward generative world m
 odels that emulate global and counterfactual reasoning\, causality\, and p
 hysics.\n\nhttps://cml.ics.uci.edu/seminars/2025-03-17-anand-bhattad
X-ALT-DESC;FMTTYPE=text/html:<html><body><b>Anand Bhattad</b>\, Research As
 sistant Professor\, Toyota Technological Institute at Chicago<br><br><b>Ti
 tle:</b> What Generative Visual Models Understand (and Don't) About the Ph
 ysical World<br><br><b>Abstract:</b> Generative visual models like Stable 
 Diffusion and Sora generate photorealistic images and videos that are near
 ly indistinguishable from real ones to a naive observer. However\, their g
 rasp of the physical world remains an open question: Do they understand 3D
  geometry\, light\, and object interactions\, or are they mere pixel parro
 ts of their training data? Through systematic probing\, I will demonstrate
  that these models surprisingly learn fundamental scene properties — int
 rinsic images such as surface normals\, depth\, albedo\, and shading — w
 ithout explicit supervision\, which enables applications like image religh
 ting. But I will also show that this knowledge is insufficient. Careful an
 alysis reveals unexpected failures: inconsistent shadows\, multiple vanish
 ing points\, and scenes that defy basic physics. All these findings sugges
 t these models excel at local texture synthesis but struggle with global r
 easoning: a crucial gap between imitation and true understanding. I will t
 hen conclude by outlining a path toward generative world models that emula
 te global and counterfactual reasoning\, causality\, and physics.<br><br><
 a href="https://cml.ics.uci.edu/seminars/2025-03-17-anand-bhattad">https:/
 /cml.ics.uci.edu/seminars/2025-03-17-anand-bhattad</a></body></html>
URL:https://cml.ics.uci.edu/seminars/2025-03-17-anand-bhattad
END:VEVENT
END:VCALENDAR
