BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171535Z
LOCATION:Hall B7 (1)\, B Block\, Level 7
DTSTART;TZID=Asia/Tokyo:20241203T134600
DTEND;TZID=Asia/Tokyo:20241203T135800
UID:siggraphasia_SIGGRAPH Asia 2024_sess105_papers_181@linklings.com
SUMMARY:Customizing Text-to-Image Diffusion with Object Viewpoint Control
DESCRIPTION:Nupur Kumari and Grace Su (Carnegie Mellon Uniersity); Richard
  Zhang, Taesung Park, and Eli Shechtman (Adobe Research); and Jun-Yan Zhu 
 (Carnegie Mellon Uniersity)\n\nModel customization introduces new concepts
  to existing text-to-image models, enabling the generation of these new co
 ncepts/objects in novel contexts.\nHowever, such methods lack accurate cam
 era view control with respect to the new object, and users must resort to 
 prompt engineering (e.g., adding "top-view'") to achieve coarse view contr
 ol. In this work, we introduce a new task -- enabling explicit control of 
 the object viewpoint in the customization of text-to-image diffusion model
 s. This allows us to modify the custom object's properties and generate it
  in various background scenes via text prompts, all while incorporating th
 e object viewpoint as an additional control. This new task presents signif
 icant challenges, as one must harmoniously merge a 3D representation from 
 the multi-view images with the 2D pre-trained model. To bridge this gap, w
 e propose to condition the diffusion process on the 3D object features ren
 dered from the target viewpoint. During training, we fine-tune the 3D feat
 ure prediction modules to reconstruct the object's appearance and geometry
 , while reducing overfitting to the input multi-view images. Our method ou
 tperforms existing image editing and model customization baselines in pres
 erving the custom object's identity while following the target object view
 point and the text prompt.\n\nRegistration Category: Full Access, Full Acc
 ess Supporter\n\nLanguage Format: English Language\n\nSession Chair: Kfir 
 Aberman (Decart AI)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_181&sess=sess105
END:VEVENT
END:VCALENDAR
