Skip to main content
← Back to seminars

Past talk · AI/ML Seminar Series

Advancing Multimodal Models Beyond Human Supervision

XuDong Wang

PhD Student, Berkeley AI Research Lab, University of California, Berkeley

Date & time
Friday, April 4, 2025 · 2:00 PM
Location
Donald Bren Hall 4011

Abstract

To advance AI toward true artificial general intelligence, it is crucial to incorporate a wider range of sensory inputs, including physical interaction, spatial navigation, and social dynamics. However, achieving the successes of self-supervised Large Language Models (LLMs) across other modalities in our physical and digital environments remains a significant challenge. In this talk, I will discuss how self-supervised learning methods can be harnessed to advance multimodal models beyond the need for human supervision. Firstly, I will highlight a series of research efforts on self-supervised visual scene understanding that leverage the capabilities of self-supervised models to segment anything without the need for 1.1 billion labeled segmentation masks. Secondly, I will demonstrate how generative and understanding models can work together synergistically. Lastly, I will explore the increasingly important techniques for learning from unlabeled or imperfect data within the context of data-centric representation learning.

About the speaker

XuDong Wang is a final-year Ph.D. student in the Berkeley AI Research (BAIR) lab at UC Berkeley, advised by Prof. Trevor Darrell, and a research scientist on the Llama Research team at GenAI, Meta. He was previously a researcher at Google DeepMind (GDM) and the International Computer Science Institute (ICSI), and a research intern at Meta's Fundamental AI Research (FAIR) labs. His research focuses on self-supervised learning, multimodal models, and machine learning. He is a recipient of the William Oldham Fellowship at UC Berkeley.