BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171532Z
LOCATION:Hall B5 (2)\, B Block\, Level 5
DTSTART;TZID=Asia/Tokyo:20241205T150800
DTEND;TZID=Asia/Tokyo:20241205T151900
UID:siggraphasia_SIGGRAPH Asia 2024_sess134_papers_485@linklings.com
SUMMARY:Lumiere: A Space-Time Diffusion Model for Video Generation
DESCRIPTION:Omer Bar-Tal (Google Research, Weizmann Institute of Science);
  Hila Chefer (Google Research, Tel Aviv University); Omer Tov, Charles Her
 rmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Am
 it Raj, Yuanzhen Li, and Michael Rubinstein (Google Research); Tomer Micha
 eli (Google Research, Technion – Israel Institute of Technology); Oliver W
 ang and Deqing Sun (Google Research); Tali Dekel (Google Research, Weizman
 n Institute of Science); and Inbar Mosseri (Google Research)\n\nWe introdu
 ce Lumiere -- a text-to-video diffusion model designed for synthesizing vi
 deos that portray realistic, diverse and coherent motion -- a pivotal chal
 lenge in video synthesis. To this end, we introduce a Space-Time U-Net arc
 hitecture that generates the entire temporal duration of the video at once
 , through a single pass in the model. This is in contrast to existing vide
 o models which synthesize distant keyframes followed by temporal super-res
 olution -- an approach that inherently makes global temporal consistency d
 ifficult to achieve. By deploying both spatial and (importantly) temporal 
 down- and up-sampling and leveraging a pre-trained text-to-image diffusion
  model, our model learns to directly generate a full-frame-rate, low-resol
 ution video by processing it in multiple space-time scales. We demonstrate
  state-of-the-art text-to-video generation results, and show that our desi
 gn easily facilitates a wide range of content creation tasks and video edi
 ting applications, including image-to-video, video inpainting, and stylize
 d generation.\n\nRegistration Category: Full Access, Full Access Supporter
 \n\nLanguage Format: English Language\n\nSession Chair: Nanxuan Zhao (Adob
 e Research)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_485&sess=sess134
END:VEVENT
END:VCALENDAR
