BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171534Z
LOCATION:Hall B5 (1)\, B Block\, Level 5
DTSTART;TZID=Asia/Tokyo:20241204T134600
DTEND;TZID=Asia/Tokyo:20241204T135800
UID:siggraphasia_SIGGRAPH Asia 2024_sess115_papers_702@linklings.com
SUMMARY:BlobGEN-3D: Compositional 3D-Consistent Freeview Image Generation 
 with 3D Blobs
DESCRIPTION:Chao Liu, Weili Nie, Sifei Liu, Abhishek Badki, Hang Su, Morte
 za Mardani, Benjamin Eckart, and Arash Vahdat (NVIDIA)\n\nRecent advances 
 in text-to-image diffusion models have significantly enhanced image genera
 tion quality, when trained on internet-scale data. However, existing metho
 ds are constrained by their reliance on image or scene-level conditions, l
 imiting their ability to synthesize composable 3D objects in a complex sce
 ne. To address these limitations, we propose BlobGEN-3D, a novel approach 
 that decouples compositional 3D scene representation from 2D image generat
 ion, enabling direct controllability in the 3D space while fully leveragin
 g the capabilities of 2D diffusion models. Specifically, BlobGEN-3D utiliz
 es object-level 3D blobs with rich textual descriptions as the 3D scene re
 presentation, which is amenable to 2D projection, and is seamlessly integr
 able with 2D diffusion models. Based on this representation, we introduce 
 an auto-regressive pipeline for freeview image generation, by conditioning
  the pretrained blob-grounded 2D text-to-image\ndiffusion model on the pre
 viously generated image. Our method has three key features: (i) it enables
  modular representation of 3D scene elements; (ii) coherent cross-view 2D 
 generation; and (iii) manipulation of object appearance in the generated i
 mage sequences. Our method not only competes with the existing multi-view 
 and optimization-based approaches, but also offers object-level appearance
  control, which was not possible before with alternatives that solely rely
  on scene-level descriptions, or image captions.\n\nRegistration Category:
  Full Access, Full Access Supporter\n\nLanguage Format: English Language\n
 \nSession Chair: Peng-Shuai Wang (Peking University)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_702&sess=sess115
END:VEVENT
END:VCALENDAR
