BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171530Z
LOCATION:Hall B5 (2)\, B Block\, Level 5
DTSTART;TZID=Asia/Tokyo:20241205T170500
DTEND;TZID=Asia/Tokyo:20241205T171600
UID:siggraphasia_SIGGRAPH Asia 2024_sess137_papers_915@linklings.com
SUMMARY:StyleCrafter: Taming Stylized Video Diffusion with Reference-Augme
 nted Adapter Learning
DESCRIPTION:Gongye Liu (Tsinghua University); Menghan Xia, Yong Zhang, and
  Haoxin Chen (Tencent AI lab); Jinbo Xing (Chinese University of Hong Kong
 ); Yibo Wang (Tsinghua University); Xintao Wang and Ying Shan (Tencent); a
 nd Yujiu Yang (Tsinghua University)\n\nText-to-video (T2V) models have sho
 wn remarkable capabilities in generating diverse videos. However, they str
 uggle to produce user-desired artistic videos due to (i) text's inherent c
 lumsiness in expressing specific styles and (ii) the generally degraded st
 yle fidelity. To address these challenges, we introduce StyleCrafter, a ge
 neric method that enhances pre-trained T2V models with a style control ada
 pter, allowing video generation in any style by feeding a reference image.
  Considering the scarcity of artistic video data, we propose to first trai
 n a style control adapter using style-rich image datasets, then transfer t
 he learned stylization ability to video generation through a tailor-made f
 inetuning paradigm. To promote content-style disentanglement, we employ ca
 refully designed data augmentation strategies to enhance decoupled learnin
 g. Additionally, we propose a scale-adaptive fusion module to balance the 
 influences of text-based content features and image-based style features, 
 which helps generalization across various text and style combinations. Sty
 leCrafter efficiently generates high-quality stylized videos that align wi
 th the content of the texts and resemble the style of the reference images
 . Experiments demonstrate that our approach is more flexible and efficient
  than existing competitors.\n\nRegistration Category: Full Access, Full Ac
 cess Supporter\n\nLanguage Format: English Language\n\nSession Chair: Mich
 ael Rubinstein (Google)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_915&sess=sess137
END:VEVENT
END:VCALENDAR
