Past talk · AI/ML Seminar Series
Foundations of Generative Discovery Beyond the Data with Flow and Diffusion Models
PhD Student, ETH AI Center
- Date & time
- Wednesday, June 10, 2026 · 1:00 PM
- Location
- Donald Bren Hall 4011
Abstract
Recent progress in flow and diffusion models has made generative models powerful priors over complex scientific design spaces, while reward-guided adaptation offers a practical way to steer them toward desired properties. This talk asks what is needed to turn such steering into discovery. I will first discuss tail-aware reward guidance: rather than maximizing average reward, one may deliberately sacrifice expected reward to concentrate probability on rare, high-value samples in the top tail of the reward distribution. I will then argue that discovery also requires debiasing the generative model itself. In the natural sciences, available data are local, limited, and biased by prior discoveries and measurement processes, so distribution matching can hide valid low-probability modes. I will present Flow Density Control (FDC) as a framework for distributional fine-tuning of flow and diffusion models, allowing to perform tasks including entropy-driven mode discovery, risk-sensitive adaptation, and experimental design. Finally, I will introduce mathematical foundations for out-of-distribution flow modeling through generable-set expansion rather than standard distribution matching. And present Active Flow Expansion (ActFlow), a synthetic pre-training method that uses verifier feedback and active exploration in learned flow representations to expand coverage over valid molecular, peptide, and protein space, while enjoying first-of-their-kind statistical guarantees for out-of-distribution generative modeling.
About the speaker
Riccardo De Santi is a PhD student and ETH AI Center Doctoral Fellow at ETH Zurich, advised by Andreas Krause, Niao He, and Kjell Jorner. He is currently visiting Caltech, working with Yisong Yue and Frances H. Arnold on generative discovery for chemistry and biology application. His research develops mathematical foundations and algorithms for discovery beyond the data, building on his earlier work on reinforcement learning exploration, which received an ICML Outstanding Paper Award, and extending these ideas to flow and diffusion models. Broadly, he aims to help build the foundations of a science of generative discovery: principled methods for discovering new, valid, and useful structures, designs, and hypotheses, with applications to molecules, peptides, and proteins.