BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171532Z
LOCATION:Hall B7 (1)\, B Block\, Level 7
DTSTART;TZID=Asia/Tokyo:20241206T153100
DTEND;TZID=Asia/Tokyo:20241206T154300
UID:siggraphasia_SIGGRAPH Asia 2024_sess150_papers_211@linklings.com
SUMMARY:PuzzleAvatar: Assembling 3D Avatars from Personal Albums
DESCRIPTION:Yuliang Xiu (Max Planck Institute for Intelligent Systems); Yu
 fei Ye (Carnegie Mellon University); Zhen Liu (Max Planck Institute for In
 telligent Systems; Mila, Université de Montréal); Dimitris Tzionas (Univer
 sity of Amsterdam); and Michael J. Black (Max Planck Institute for Intelli
 gent Systems)\n\nGenerating personalized 3D avatars is crucial for AR/VR. 
 However, recent text-to-3D methods that generate avatars for celebrities o
 r fictional characters, struggle with everyday people. Methods for faithfu
 l reconstruction typically require full-body images in controlled settings
 . What if a user could just upload their personal "OOTD" (Outfit Of The Da
 y) photo collection and get a faithful avatar in return? The challenge is 
 that such casual photo collections contain diverse poses, challenging view
 points, cropped views, and occlusion (albeit with a consistent outfit, acc
 essories and hairstyle). We address this novel "Album2Human" task by devel
 oping PuzzleAvatar, a novel model that generates a faithful 3D avatar (in 
 a canonical pose) from a personal OOTD album, while bypassing the challeng
 ing estimation of body and camera pose. To this end, we fine-tune a founda
 tional vision-language model (VLM) on such photos, encoding the appearance
 , identity, garments, hairstyles, and accessories of a person into (separa
 te) learned tokens and instilling these cues into the VLM. In effect, we e
 xploit the learned tokens as "puzzle pieces" from which we assemble a fait
 hful, personalized 3D avatar. Importantly, we can customize avatars by sim
 ply inter-changing tokens. As a benchmark for this new task, we collect a 
 new dataset, called PuzzleIOI, with 41 subjects in a total of nearly 1K OO
 TD configurations, in challenging partial photos with paired ground-truth 
 3D bodies. Evaluation shows that PuzzleAvatar not only has high reconstruc
 tion accuracy, outperforming TeCH and MVDreamBooth, but also a unique scal
 ability to album photos, and strong robustness. Our model and data will be
  public.\n\nRegistration Category: Full Access, Full Access Supporter\n\nL
 anguage Format: English Language\n\nSession Chair: Li-Yi Wei (Adobe Resear
 ch)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_211&sess=sess150
END:VEVENT
END:VCALENDAR
