BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//UC Irvine//CML Seminars//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Enabling Language Models to Process Information at Scale
X-WR-TIMEZONE:America/Los_Angeles
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:20070311T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:20071104T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
UID:2025-03-13-tianyu-gao@cml.ics.uci.edu
DTSTAMP:20250313T000000Z
SEQUENCE:57690
DTSTART;TZID=America/Los_Angeles:20250313T110000
DTEND;TZID=America/Los_Angeles:20250313T120000
SUMMARY:[CML Seminar] Tianyu Gao: Enabling Language Models to Process Infor
 mation at Scale
LOCATION:Donald Bren Hall 4011
DESCRIPTION:Tianyu Gao\, PhD Student\, Department of Computer Science\, Pri
 nceton University\n\nTitle: Enabling Language Models to Process Informatio
 n at Scale\n\nAbstract: Language models (LMs) can effectively internalize 
 knowledge from vast amounts of pre-training data\, enabling them to achiev
 e remarkable performance on exam-style benchmarks. Expanding their ability
  to compile\, synthesize\, and reason over large volumes of information on
  the fly will further unlock transformative applications\, ranging from AI
  literature assistants to generative search engines. In this talk\, I will
  present my research on advancing LMs for processing information at scale.
  (1) I will present my evaluation framework for LM-based information-seeki
 ng systems\, emphasizing the importance of providing citations for verifyi
 ng the model-generated answers. (2) I will then introduce my foundational 
 work on using contrastive learning to produce high-performing text embeddi
 ngs\, which form the cornerstone of effective and scalable search. (3) In 
 addition to building systems that can process large-scale information\, I 
 will discuss my contributions to creating efficient pre-training and custo
 mization methods for LMs. Finally\, I will share my vision for the next ge
 neration of autonomous information processing systems.\n\nhttps://cml.ics.
 uci.edu/seminars/2025-03-13-tianyu-gao
X-ALT-DESC;FMTTYPE=text/html:<html><body><b>Tianyu Gao</b>\, PhD Student\, 
 Department of Computer Science\, Princeton University<br><br><b>Title:</b>
  Enabling Language Models to Process Information at Scale<br><br><b>Abstra
 ct:</b> Language models (LMs) can effectively internalize knowledge from v
 ast amounts of pre-training data\, enabling them to achieve remarkable per
 formance on exam-style benchmarks. Expanding their ability to compile\, sy
 nthesize\, and reason over large volumes of information on the fly will fu
 rther unlock transformative applications\, ranging from AI literature assi
 stants to generative search engines. In this talk\, I will present my rese
 arch on advancing LMs for processing information at scale. (1) I will pres
 ent my evaluation framework for LM-based information-seeking systems\, emp
 hasizing the importance of providing citations for verifying the model-gen
 erated answers. (2) I will then introduce my foundational work on using co
 ntrastive learning to produce high-performing text embeddings\, which form
  the cornerstone of effective and scalable search. (3) In addition to buil
 ding systems that can process large-scale information\, I will discuss my 
 contributions to creating efficient pre-training and customization methods
  for LMs. Finally\, I will share my vision for the next generation of auto
 nomous information processing systems.<br><br><a href="https://cml.ics.uci
 .edu/seminars/2025-03-13-tianyu-gao">https://cml.ics.uci.edu/seminars/2025
 -03-13-tianyu-gao</a></body></html>
URL:https://cml.ics.uci.edu/seminars/2025-03-13-tianyu-gao
END:VEVENT
END:VCALENDAR
