BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171531Z
LOCATION:Hall B5 (2)\, B Block\, Level 5
DTSTART;TZID=Asia/Tokyo:20241204T104500
DTEND;TZID=Asia/Tokyo:20241204T105600
UID:siggraphasia_SIGGRAPH Asia 2024_sess113_papers_683@linklings.com
SUMMARY:Quark: Real-time, High-resolution, and General Neural View Synthes
 is
DESCRIPTION:John Flynn, Michael Broxton, Lukas Murmann, Lucy Chai, Matthew
  DuVall, Clément Godard, Kathryn Heal, Srinivas Kaza, Stephen Lombardi, Xu
 an Luo, Supreeth Achar, Kira Prabhu, Tiancheng Sun, Lynn Tsai, and Ryan Ov
 erbeck (Google)\n\nWe present a novel neural algorithm for performing high
 -quality, high-resolution, real-time novel view synthesis. From a sparse s
 et of input RGB images or videos streams, our network both reconstructs th
 e 3D scene and renders novel views at 1080p resolution at 30fps on an NVID
 IA A100. Our feed-forward network generalizes across a wide variety of dat
 asets and scenes and produces state-of-the-art quality for a real-time met
 hod. Our quality approaches, and in some cases surpasses, the quality of s
 ome of the top offline methods. In order to achieve these results we use a
  novel combination of several key concepts, and tie them together into a c
 ohesive and effective algorithm. We build on previous works that represent
  the scene using semi-transparent layers and use an iterative learned rend
 er-and-refine approach to improve those layers. Instead of flat layers, ou
 r method reconstructs layered depth maps (LDMs) that efficiently represent
  scenes with complex depth and occlusions. The iterative update steps are 
 embedded in a multi-scale, UNet-style architecture to perform as much comp
 ute as possible at reduced resolution. Within each update step, to better 
 aggregate the information from multiple input views, we use a specialized 
 Transformer-based network component. This allows the majority of the per-i
 nput image processing to be performed in the input image space, as opposed
  to layer space, further increasing efficiency. Finally, due to the real-t
 ime nature of our reconstruction and rendering, we dynamically create and 
 discard the internal 3D geometry for each frame, generating the LDM for ea
 ch view. Taken together, this produces a novel and effective algorithm for
  view synthesis. Through extensive evaluation, we demonstrate that we achi
 eve state-of-the-art quality at real-time rates.\n\nRegistration Category:
  Full Access, Full Access Supporter\n\nLanguage Format: English Language\n
 \nSession Chair: Forrester Cole (Google)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_683&sess=sess113
END:VEVENT
END:VCALENDAR
