BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171532Z
LOCATION:Hall B5 (2)\, B Block\, Level 5
DTSTART;TZID=Asia/Tokyo:20241204T172800
DTEND;TZID=Asia/Tokyo:20241204T174000
UID:siggraphasia_SIGGRAPH Asia 2024_sess122_papers_667@linklings.com
SUMMARY:Camera Settings as Tokens: Modeling Photography on Latent Diffusio
 n Models
DESCRIPTION:I-Sheng Fang, Yue-Hua Han, and Jun-Cheng Chen (Academia Sinica
 )\n\nText-to-image models have revolutionized content creation, enabling u
 sers to generate images from natural language prompts. While recent advanc
 ements in conditioning these models offer more control over the generated 
 results, photography—a significant artistic domain—remains inadequately in
 tegrated into these systems. Our research identifies critical gaps in mode
 ling camera settings and photographic terms within text-to-image synthesis
 . Vision-language models (VLMs) like CLIP and OpenCLIP, which typically dr
 ive the text conditions through cross-attention mechanisms of conditional 
 diffusion models, struggle to represent numerical data like camera setting
 s effectively in their textual space. To address these challenges, we pres
 ent CameraSettings20k, a new dataset aggregated from RAISE, DDPD, and PPR1
 0K.Our curated dataset offers normalized camera settings for over 20,000 r
 aw-format images, providing equivalent values standardized to a full-frame
  sensor. Furthermore, we introduce Camera Settings as Tokens, an embedding
  approach leveraging the LoRA adapter of Latent Diffusion Models (LDMs) to
  numerically control image generation based on photographic principles lik
 e focal length, aperture, film speed, and exposure time. Our experimental 
 results demonstrate the effectiveness of the proposed approach to generate
  promising synthesized images obeying the photographic principles given th
 e specified numerical camera settings. Furthermore, our work not only brid
 ges the gap between camera settings and user-friendly photographic control
  in image synthesis but also sets the stage for future explorations into m
 ore physics-aware generative models.\n\nRegistration Category: Full Access
 , Full Access Supporter\n\nLanguage Format: English Language\n\nSession Ch
 air: Minhyuk Sung (Korea Advanced Institute of Science and Technology (KAI
 ST))\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_667&sess=sess122
END:VEVENT
END:VCALENDAR
