BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171536Z
LOCATION:Hall B5 (2)\, B Block\, Level 5
DTSTART;TZID=Asia/Tokyo:20241204T131100
DTEND;TZID=Asia/Tokyo:20241204T132300
UID:siggraphasia_SIGGRAPH Asia 2024_sess116_papers_884@linklings.com
SUMMARY:InstantDrag: Improving Interactivity in Drag-based Image Editing
DESCRIPTION:Joonghyuk Shin (Seoul National University), Daehyeon Choi (POS
 TECH), and Jaesik Park (Seoul National University)\n\nDrag-based image edi
 ting has recently gained popularity for its interactivity and precision. H
 owever, despite the ability of text-to-image models to generate samples wi
 thin a second, drag editing still lags behind due to the challenge of accu
 rately reflecting user interaction while maintaining image content. Some e
 xisting approaches rely on computationally intensive per-image optimizatio
 n or intricate guidance-based methods, requiring additional inputs such as
  masks for movable regions and text prompts, thereby compromising the inte
 ractivity of the editing process. We introduce InstantDrag, an optimizatio
 n-free pipeline that enhances interactivity and speed, requiring only an i
 mage and a drag instruction as input. InstantDrag consists of two carefull
 y designed networks: a drag-conditioned optical flow generator (FlowGen) a
 nd an optical flow-conditioned diffusion model (FlowDiffusion). InstantDra
 g learns motion dynamics for drag-based image editing in real-world video 
 datasets by decomposing the task into motion generation and motion-conditi
 oned image generation. We demonstrate InstantDrag's capability to perform 
 fast, photo-realistic edits without masks or text prompts through experime
 nts on facial video datasets and general scenes. These results highlight t
 he efficiency of our approach in handling drag-based image editing, making
  it a promising solution for interactive, real-time applications.\n\nRegis
 tration Category: Full Access, Full Access Supporter\n\nLanguage Format: E
 nglish Language\n\nSession Chair: Dani Lischinski (Hebrew University of Je
 rusalem, Google)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_884&sess=sess116
END:VEVENT
END:VCALENDAR
