BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:Asia/Tokyo
X-LIC-LOCATION:Asia/Tokyo
BEGIN:STANDARD
TZOFFSETFROM:+0900
TZOFFSETTO:+0900
TZNAME:JST
DTSTART:18871231T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260817T171531Z
LOCATION:Hall B7 (1)\, B Block\, Level 7
DTSTART;TZID=Asia/Tokyo:20241206T134200
DTEND;TZID=Asia/Tokyo:20241206T135600
UID:siggraphasia_SIGGRAPH Asia 2024_sess147_papers_168@linklings.com
SUMMARY:Dance-to-Music Generation with Encoder-based Textual Inversion
DESCRIPTION:Sifei Li, Weiming Dong, and Yuxin Zhang (MAIS, Institute of Au
 tomation, Chinese Academy of Sciences; School of Artificial Intelligence, 
 University of Chinese Academy of Sciences); Fan Tang (University of Chines
 e Academy of Sciences); Chongyang Ma (Kuaishou Technology); Oliver Deussen
  (University of Konstanz); Tong-Yee Lee (National Cheng-Kung University); 
 and Changsheng Xu (MAIS, Institute of Automation, Chinese Academy of Scien
 ces; School of Artificial Intelligence, University of Chinese Academy of S
 ciences)\n\nThe seamless integration of music with dance movements is esse
 ntial for communicating the artistic intent of a dance piece. This alignme
 nt also significantly improves the immersive quality of gaming experiences
  and animation productions. Although there has been remarkable advancement
  in creating high-fidelity music from textual descriptions, current method
 ologies mainly focus on modulating overall characteristics such as genre a
 nd emotional tone. They often overlook the nuanced management of temporal 
 rhythm, which is indispensable in crafting music for dance, since it intri
 cately aligns the musical beats with the dancers' movements. Recognizing t
 his gap, we propose an encoder-based textual inversion technique to augmen
 t text-to-music models with visual control, facilitating personalized musi
 c generation. Specifically, we develop dual-path rhythm-genre inversion to
  effectively integrate the rhythm and genre of a dance motion sequence int
 o the textual space of a text-to-music model. Contrary to traditional text
 ual inversion methods, which directly update text embeddings to reconstruc
 t a single target object, our approach utilizes separate rhythm and genre 
 encoders to obtain text embeddings for two pseudo-words, adapting to the v
 arying rhythms and genres. We collect a new dataset called In-the-wild Dan
 ce Videos (InDV) and demonstrate that our approach outperforms state-of-th
 e-art methods across multiple evaluation metrics. Furthermore, our method 
 is able to adapt to changes in tempo} and effectively integrates with the 
 inherent text-guided generation capability of the pre-trained model. Our s
 ource code and demo videos are available at https://github.com/lsfhuihuiff
 /Dance-to-music_Siggraph_Asia_2024.\n\nRegistration Category: Full Access,
  Full Access Supporter\n\nLanguage Format: English Language\n\nSession Cha
 ir: Yi Zhou (Roblox)\n\n
URL:https://asia.siggraph.org/2024/program/?id=papers_168&sess=sess147
END:VEVENT
END:VCALENDAR
