BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171533Z
LOCATION:Hall B5 (1)\, B Block\, Level 5
DTSTART;TZID=Asia/Tokyo:20241206T091100
DTEND;TZID=Asia/Tokyo:20241206T092300
UID:siggraphasia_SIGGRAPH Asia 2024_sess139_papers_1298@linklings.com
SUMMARY:SPARK: Self-supervised Personalized Real-time Monocular Face Captu
 re
DESCRIPTION:Kelian Baert (Technicolor Group, Institut national de recherch
 e en informatique et en automatique (INRIA) Rennes); Shrisha Bharadwaj (Ma
 x Planck Institute for Intelligent Systems); Fabien Castan and Benoit Mauj
 ean (Technicolor Group); Marc Christie (Institut national de recherche en 
 informatique et en automatique (INRIA)); Victoria Fernández Abrevaya (Max 
 Planck Institute for Intelligent Systems); and Adnane Boukhayma (Institut 
 national de recherche en informatique et en automatique (INRIA))\n\nFeedfo
 rward monocular face capture methods seek to reconstruct posed faces from 
 a single image of a person. Current state of the art approaches have the a
 bility to regress parametric 3D face models in real-time across a wide ran
 ge of identities, lighting conditions and poses by leveraging large image 
 datasets of human faces. These methods however suffer from clear limitatio
 ns in that the underlying parametric face model only provides a coarse est
 imation of the face shape, thereby limiting their practical applicability 
 in tasks that require precise 3D reconstruction (aging, face swapping, dig
 ital make-up,...).\n\nIn this paper, we propose a method for high-precisio
 n 3D face capture taking advantage of a collection of unconstrained videos
  of a subject as prior information.  Our proposal builds on a two stage ap
 proach. We start with the reconstruction of a detailed 3D face avatar of t
 he person, capturing both precise geometry and appearance from a collectio
 n of videos. We then use the encoder from a pre-trained monocular face rec
 onstruction method, substituting its decoder with our personalized model, 
 and proceed with transfer learning on the video collection. Using our pre-
 estimated image formation model, we obtain a more precise self-supervision
  objective, enabling improved expression and pose alignment. This results 
 in a trained encoder capable of efficiently regressing pose and expression
  parameters in real-time from previously unseen images, which combined wit
 h our personalized geometry model yields more accurate and high fidelity  
 mesh inference.   \n    \nThrough extensive qualitative and quantitative e
 valuation, we showcase the superiority of our final model as compared to s
 tate-of-the-art baselines, and demonstrate its generalization ability to u
 nseen pose, expression and lighting.\n\nRegistration Category: Full Access
 , Full Access Supporter\n\nLanguage Format: English Language\n\nSession Ch
 air: Kui Wu (LIGHTSPEED)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_1298&sess=sess139
END:VEVENT
END:VCALENDAR
