English
Related papers

Related papers: Text2Move: Text-to-moving sound generation via tra…

200 papers

World models have demonstrated impressive performance on robotic learning tasks. Many such tasks inherently demand multimodal reasoning; for example, filling a bottle with water will lead to visual information alone being ambiguous or…

Robotics · Computer Science 2025-12-10 Fan Zhang , Michael Gienger

Text-to-motion generation is a crucial task in computer vision, which generates the target 3D motion by the given text. The existing annotated datasets are limited in scale, resulting in most existing methods overfitting to the small…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Ke Fan , Jiangning Zhang , Ran Yi , Jingyu Gong , Yabiao Wang , Yating Wang , Xin Tan , Chengjie Wang , Lizhuang Ma

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

In this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the first to explore the combination of pre-trained text-to-audio…

Developing algorithms for sound classification, detection, and localization requires large amounts of flexible and realistic audio data, especially when leveraging modern machine learning and beamforming techniques. However, most existing…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Luca Barbisan , Marco Levorato , Fabrizio Riente

We present a new method for text-driven motion transfer - synthesizing a video that complies with an input text prompt describing the target objects and scene while maintaining an input video's motion and scene layout. Prior methods are…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Danah Yatim , Rafail Fridman , Omer Bar-Tal , Yoni Kasten , Tali Dekel

For immersive applications, the generation of binaural sound that matches its visual counterpart is crucial to bring meaningful experiences to people in a virtual environment. Recent studies have shown the possibility of using neural…

Sound · Computer Science 2023-05-22 Francesc Lluís , Vasileios Chatziioannou , Alex Hofmann

Text-Image-to-Video (TI2V) generation aims to generate a video from an image following a text description, which is also referred to as text-guided image animation. Most existing methods struggle to generate videos that align well with the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Shijie Wang , Samaneh Azadi , Rohit Girdhar , Saketh Rambhatla , Chen Sun , Xi Yin

Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to reduce the amount of unexplained variation in training data…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Devang S Ram Mohan , Vivian Hu , Tian Huey Teh , Alexandra Torresquintero , Christopher G. R. Wallis , Marlene Staib , Lorenzo Foglianti , Jiameng Gao , Simon King

Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR,…

Recent work on audio-visual navigation assumes a constantly-sounding target and restricts the role of audio to signaling the target's position. We introduce semantic audio-visual navigation, where objects in the environment make sounds…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Changan Chen , Ziad Al-Halah , Kristen Grauman

Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, producing both modalities simultaneously still typically relies on…

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models to learn the…

Sound · Computer Science 2024-03-14 Shentong Mo , Jing Shi , Yapeng Tian

Recent advances in diffusion models have showcased promising results in the text-to-video (T2V) synthesis task. However, as these T2V models solely employ text as the guidance, they tend to struggle in modeling detailed temporal dynamics.…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Seungwoo Lee , Chaerin Kong , Donghyeon Jeon , Nojun Kwak

Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emotional nuance, dynamic prosody, and lyric-based semantics,…

Graphics · Computer Science 2025-09-03 Zikai Huang , Yihan Zhou , Xuemiao Xu , Cheng Xu , Xiaofen Xing , Jing Qin , Shengfeng He

While most music generation models use textual or parametric conditioning (e.g. tempo, harmony, musical genre), we propose to condition a language model based music generation system with audio input. Our exploration involves two distinct…

Sound · Computer Science 2024-07-31 Simon Rouard , Yossi Adi , Jade Copet , Axel Roebel , Alexandre Défossez

Sound is an information-rich medium that captures dynamic physical events. This work presents STReSSD, a framework that uses sound to bridge the simulation-to-reality gap for stochastic dynamics, demonstrated for the canonical case of a…

Robotics · Computer Science 2020-11-09 Carolyn Matl , Yashraj Narang , Dieter Fox , Ruzena Bajcsy , Fabio Ramos

Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene Synthesis challenge held as part of the Detection and…

In the near future, more and more machines will perform tasks in the vicinity of human spaces or support them directly in their spatially bound activities. In order to simplify the verbal communication and the interaction between robotic…

Machine Learning · Computer Science 2020-04-14 Sebastian Feld , Steffen Illium , Andreas Sedlmeier , Lenz Belzner

Synthesizing natural head motion to accompany speech for an embodied conversational agent is necessary for providing a rich interactive experience. Most prior works assess the quality of generated head motion by comparing them against a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-27 Trisha Mittal , Zakaria Aldeneh , Masha Fedzechkina , Anurag Ranjan , Barry-John Theobald