BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171531Z
LOCATION:Hall B7 (1)\, B Block\, Level 7
DTSTART;TZID=Asia/Tokyo:20241205T163000
DTEND;TZID=Asia/Tokyo:20241205T164100
UID:siggraphasia_SIGGRAPH Asia 2024_sess138_papers_912@linklings.com
SUMMARY:VOODOO XP: Expressive One-Shot Head Reenactment for VR Telepresenc
 e
DESCRIPTION:Phong Tran (MBZUAI); Egor Zakharov (ETH Zurich); Long-Nhat Ho,
  Adilbek Karmanov, and Ariana Bermudez Venegas (MBZUAI); McLean Goldwhite,
  Aviral Agarwal, and Liwen Hu (Pinscreen); Anh Tran (VinAI Research); and 
 Hao Li (MBZUAI, Pinscreen)\n\nWe introduce VOODOO XP: a 3D-aware one-shot 
 head reenactment method that can generate highly expressive facial express
 ions from any input driver video and a single 2D portrait. Our solution is
  real-time, view-consistent, and can be instantly used without calibration
  or fine-tuning. We demonstrate our solution on a monocular video setting 
 and an end-to-end VR telepresence system for two-way communication. Compar
 ed to 2D head reenactment methods, 3D-aware approaches aim to preserve the
  identity of the subject and ensure view-consistent facial geometry for no
 vel camera poses, which makes them suitable for immersive applications. Wh
 ile various facial disentanglement techniques have been introduced, cuttin
 g-edge 3D-aware neural reenactment techniques still lack expressiveness an
 d fail to reproduce complex and fine-scale facial expressions. We present 
 a novel cross-reenactment architecture that directly transfers the driver'
 s facial expressions to transformer blocks of the input source's 3D liftin
 g module. We show that highly effective disentanglement is possible using 
 an innovative multi-stage self-supervision approach, which is based on a c
 oarse-to-fine strategy, combined with an explicit face neutralization and 
 3D lifted frontalization during its initial training stage. We further int
 egrate our novel head reenactment solution into an accessible high-fidelit
 y VR telepresence system, where any person can instantly build a personali
 zed neural head avatar from any photo and bring it to life using the heads
 et. We demonstrate state-of-the-art performance in terms of expressiveness
  and likeness preservation on a large set of diverse subjects and capture 
 conditions.\n\nRegistration Category: Full Access, Full Access Supporter\n
 \nLanguage Format: English Language\n\nSession Chair: Hongbo Fu (Hong Kong
  University of Science and Technology)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_912&sess=sess138
END:VEVENT
END:VCALENDAR
