BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171530Z
LOCATION:Hall B5 (2)\, B Block\, Level 5
DTSTART;TZID=Asia/Tokyo:20241204T110800
DTEND;TZID=Asia/Tokyo:20241204T111900
UID:siggraphasia_SIGGRAPH Asia 2024_sess113_papers_253@linklings.com
SUMMARY:Cafca: High-quality Novel View Synthesis of Expressive Faces from 
 Casual Few-shot Captures
DESCRIPTION:Marcel C. Buehler and Gengyan Li (ETH Zürich, Google VR); Erro
 ll Wood, Leonhard Helminger, Xu Chen, Tanmay Shah, Daoye Wang, Stephan Gar
 bin, and Sergio Orts Escolano (Google VR); Otmar Hilliges (ETH Zürich); an
 d Dmitry Lagun, Jérémy Riviere, Paulo Gotardo, Thabo Beeler, Abhimitra Mek
 a, and Kripasindhu Sarkar (Google VR)\n\nVolumetric modeling and neural ra
 diance field representations have revolutionized 3D face capture and photo
 realistic novel view synthesis. However, these methods often require hundr
 eds of multi-view input images and are thus inapplicable to cases with les
 s than a handful of inputs.\nWe present a novel volumetric prior on human 
 faces that allows for high-fidelity expressive face modeling from as few a
 s three input views captured in the wild. Our key insight is that an impli
 cit prior trained on synthetic data alone can generalize to extremely chal
 lenging real-world identities and expressions and render novel views with 
 fine idiosyncratic details like wrinkles and eyelashes.\nWe leverage a 3D 
 Morphable Face Model to synthesize a large training set, rendering each id
 entity with different expressions, hair, clothing, and other assets. We th
 en train a conditional Neural Radiance Field prior on this synthetic datas
 et and, at inference time, fine-tune the model on a very sparse set of rea
 l images of a single subject. On average, the fine-tuning requires only th
 ree inputs to cross the synthetic-to-real domain gap. The resulting person
 alized 3D model reconstructs strong idiosyncratic facial expressions and o
 utperforms the state-of-the-art in high-quality novel view synthesis of fa
 ces from sparse inputs in terms of perceptual and photo-metric quality.\n\
 nRegistration Category: Full Access, Full Access Supporter\n\nLanguage For
 mat: English Language\n\nSession Chair: Forrester Cole (Google)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_253&sess=sess113
END:VEVENT
END:VCALENDAR
