BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171531Z
LOCATION:Hall B7 (1)\, B Block\, Level 7
DTSTART;TZID=Asia/Tokyo:20241205T170500
DTEND;TZID=Asia/Tokyo:20241205T171600
UID:siggraphasia_SIGGRAPH Asia 2024_sess138_papers_215@linklings.com
SUMMARY:TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenac
 tment with Diffusion Model
DESCRIPTION:Jiazhi Guan (Tsinghua University); Quanwei Yang (University of
  Science and Technology of China); Kaisiyuan Wang, Hang Zhou, Shengyi He, 
 Zhiliang Xu, Haocheng Feng, Errui Ding, and Jingdong Wang (Baidu); Hongtao
  Xie (University of Science and Technology of China); Youjian Zhao (Tsingh
 ua University); and Ziwei Liu (Nanyang Technological University (NTU))\n\n
 Recently, 2D speaking avatars have increasingly participated in everyday s
 cenarios due to the fast development of facial animation techniques. Howev
 er, most existing works neglect the explicit control of human bodies. In t
 his paper, we propose to drive not only the faces but also the torso and g
 esture movements of a speaking figure. Inspired by recent advances in diff
 usion models, we propose the Motion-Enhanced Textural-Aware ModeLing for S
 peaKing Avatar Reenactment (TALK-Act) framework, which enables high-fideli
 ty avatar reenactment from only short footage of monocular video. Our key 
 idea is to enhance the textural awareness with explicit motion guidance in
  diffusion modeling. Specifically, we carefully construct 2D and 3D struct
 ural information as intermediate guidance. While recent diffusion models a
 dopt a side network for control information injection, they fail to synthe
 size temporally stable results even with person-specific fine-tuning. We p
 ropose a Motion-Enhanced Textural Alignment module to enhance the bond bet
 ween driving and target signals. Moreover, we build a Memory-based Hand-Re
 covering module to help with the difficulties in hand-shape preserving. Af
 ter pre-training, our model can achieve high-fidelity 2D avatar reenactmen
 t with only 30 seconds of person-specific data. Extensive experiments demo
 nstrate the effectiveness and superiority of our proposed framework.\n\nRe
 gistration Category: Full Access, Full Access Supporter\n\nLanguage Format
 : English Language\n\nSession Chair: Hongbo Fu (Hong Kong University of Sc
 ience and Technology)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_215&sess=sess138
END:VEVENT
END:VCALENDAR
