BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171534Z
LOCATION:Hall B7 (1)\, B Block\, Level 7
DTSTART;TZID=Asia/Tokyo:20241205T165300
DTEND;TZID=Asia/Tokyo:20241205T170500
UID:siggraphasia_SIGGRAPH Asia 2024_sess138_papers_517@linklings.com
SUMMARY:PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
DESCRIPTION:Longhao Zhang, Shuang Liang, Zhipeng Ge, and Tianshu Hu (Byted
 ance)\n\nFor audio-driven visual dubbing, it remains a considerable challe
 nge to uphold and highlight speaker's persona while synthesizing accurate 
 lip synchronization. Existing methods fall short of capturing speaker's un
 ique speaking style or preserving facial details. In this paper, we presen
 t PersonaTalk, an attention-based two-stage framework, including geometry 
 construction and face rendering, for high-fidelity and personalized visual
  dubbing. In the first stage, we propose a style-aware audio encoding modu
 le that injects speaking style into audio features through a cross-attenti
 on layer. The stylized audio features are then used to drive speaker's tem
 plate geometry to obtain lip-synced geometries. In the second stage, a dua
 l-attention face renderer is introduced to render textures for the target 
 geometries. It consists of two parallel cross-attention layers, namely lip
 -attention and face-attention, which respectively sample textures from dif
 ferent reference frames to render the entire face. With our innovative des
 ign, intricate facial details can be well preserved. Comprehensive experim
 ents and user studies demonstrate our advantages over other state-of-the-a
 rt methods in terms of visual quality, lip-sync accuracy and persona prese
 rvation. Furthermore, as a person-generic framework, PersonaTalk can achie
 ve competitive performance as state-of-the-art person-specific methods.\n\
 nRegistration Category: Full Access, Full Access Supporter\n\nLanguage For
 mat: English Language\n\nSession Chair: Hongbo Fu (Hong Kong University of
  Science and Technology)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_517&sess=sess138
END:VEVENT
END:VCALENDAR
