BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171531Z
LOCATION:Hall B7 (1)\, B Block\, Level 7
DTSTART;TZID=Asia/Tokyo:20241206T130000
DTEND;TZID=Asia/Tokyo:20241206T131400
UID:siggraphasia_SIGGRAPH Asia 2024_sess147_tog_106@linklings.com
SUMMARY:Speed-Aware Audio-Driven Speech Animation using Adaptive Windows
DESCRIPTION:Sunjin Jung (KAIST, Visual Media Lab); Yeongho Seol (NVIDIA); 
 Kwanggyoon Seo and Hyeonho Na (KAIST, Visual Media Lab); Seonghyeon Kim (K
 AIST, Visual Media Lab; Anigma Technologies); and Vanessa Tan and Junyong 
 Noh (KAIST, Visual Media Lab)\n\nWe present a novel method that can genera
 te realistic speech animations of a 3D face from audio using multiple adap
 tive windows. In contrast to previous studies that use a fixed size audio 
 window, our method accepts an adaptive audio window as input, reflecting t
 he audio speaking rate to use consistent phonemic information. Our system 
 consists of three parts. First, the speaking rate is estimated from the in
 put audio using a neural network trained in a self-supervised manner. Seco
 nd, the appropriate window size that encloses the audio features is predic
 ted adaptively based on the estimated speaking rate. Another key element l
 ies in the use of multiple audio windows of different sizes as input to th
 e animation generator: a small window to concentrate on detailed informati
 on and a large window to consider broad phonemic information near the cent
 er frame. Finally, the speech animation is generated from the multiple ada
 ptive audio windows. Our method can generate realistic speech animations f
 rom in-the-wild audios at any speaking rate, i.e., fast raps, slow songs, 
 as well as normal speech. We demonstrate via extensive quantitative and qu
 alitative evaluations including a user study that our method outperforms s
 tate-of-the-art approaches.\n\nRegistration Category: Full Access, Full Ac
 cess Supporter\n\nLanguage Format: English Language\n\nSession Chair: Yi Z
 hou (Roblox)\n\n
URL:https://asia.siggraph.org/2024/program/?id=tog_106&sess=sess147
END:VEVENT
END:VCALENDAR
