BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171532Z
LOCATION:Hall B5 (2)\, B Block\, Level 5
DTSTART;TZID=Asia/Tokyo:20241205T144500
DTEND;TZID=Asia/Tokyo:20241205T145600
UID:siggraphasia_SIGGRAPH Asia 2024_sess134_papers_608@linklings.com
SUMMARY:Still-Moving: Customized Video Generation without Customized Video
  Data
DESCRIPTION:Hila Chefer (Google Research, Tel Aviv University); Shiran Zad
 a, Roni Paiss, Ariel Ephrat, Omer Tov, and Michael Rubinstein (Google Rese
 arch); Lior Wolf (Tel Aviv University); Tali Dekel (Google Research, Weizm
 ann Institute of Science); Tomer Michaeli (Google Research, Technion – Isr
 ael Institute of Technology); and Inbar Mosseri (Google Research)\n\nCusto
 mizing text-to-image (T2I) models has seen tremendous progress recently, p
 articularly in areas such as personalization, stylization, and conditional
  generation. However, expanding this progress to video generation is still
  in its infancy, primarily due to the lack of customized video data. \nIn 
 this work, we introduce Still-Moving, a novel generic framework for custom
 izing a text-to-video (T2V) model, without requiring any customized video 
 data. The framework applies to the prominent T2V design where the video mo
 del is built over a text-to-image (T2I) model (e.g., via inflation). We as
 sume access to a customized version of the T2I model, trained only on stil
 l image data (e.g., using DreamBooth or StyleDrop).\nNaively plugging in t
 he weights of the customized T2I model into the T2V model often leads to s
 ignificant artifacts or insufficient adherence to the customization data. 
 \nTo overcome this issue, we train lightweight Spatial Adapters that adjus
 t the features produced by the injected T2I layers.\nImportantly, our adap
 ters are trained on "frozen videos" (i.e., repeated images), constructed f
 rom image samples generated by the customized T2I model. This training is 
 facilitated by a novel Motion Adapter module, which allows us to train on 
 such static videos while preserving the motion prior of the video model. A
 t test time, we remove the Motion Adapter modules and leave in only the tr
 ained Spatial Adapters. This restores the motion prior of the T2V model wh
 ile adhering to the spatial prior of the customized T2I model.\nWe demonst
 rate the effectiveness of our approach on diverse tasks including personal
 ized, stylized, and conditional generation. In all evaluated scenarios, ou
 r method seamlessly integrates the spatial prior of the customized T2I mod
 el with a motion prior supplied by the T2V model.\n\nRegistration Category
 : Full Access, Full Access Supporter\n\nLanguage Format: English Language\
 n\nSession Chair: Nanxuan Zhao (Adobe Research)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_608&sess=sess134
END:VEVENT
END:VCALENDAR
