English
Related papers

Related papers: MOSPA: Human Motion Generation Driven by Spatial A…

200 papers

The body movements accompanying speech aid speakers in expressing their ideas. Co-speech motion generation is one of the important approaches for synthesizing realistic avatars. Due to the intricate correspondence between speech and motion,…

Multimedia · Computer Science 2024-08-28 Sen Wang , Jiangning Zhang , Xin Tan , Zhifeng Xie , Chengjie Wang , Lizhuang Ma

The accompanying actions and gestures in dialogue are often closely linked to interactions with the environment, such as looking toward the interlocutor or using gestures to point to the described target at appropriate moments. Speech and…

Human motion modeling is important for many modern graphics applications, which typically require professional skills. In order to remove the skill barriers for laymen, recent motion generation methods can directly generate human motions…

Computer Vision and Pattern Recognition · Computer Science 2022-09-01 Mingyuan Zhang , Zhongang Cai , Liang Pan , Fangzhou Hong , Xinying Guo , Lei Yang , Ziwei Liu

The recent Segment Anything Model 2 (SAM2) has demonstrated exceptional capabilities in interactive object segmentation for both images and videos. However, as a foundational model on interactive segmentation, SAM2 performs segmentation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Qiushi Yang , Yuan Yao , Miaomiao Cui , Liefeng Bo

Learning human motion based on a time-dependent input signal presents a challenging yet impactful task with various applications. The goal of this task is to generate or estimate human movement that consistently reflects the temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Quang Nguyen , Tri Le , Baoru Huang , Minh Nhat Vu , Ngan Le , Thieu Vo , Anh Nguyen

We introduce an approach to convert mono audio recorded by a 360 video camera into spatial audio, a representation of the distribution of sound over the full viewing sphere. Spatial audio is an important component of immersive 360 video…

Sound · Computer Science 2018-09-10 Pedro Morgado , Nuno Vasconcelos , Timothy Langlois , Oliver Wang

Generating reasonable and high-quality human interactive motions in a given dynamic environment is crucial for understanding, modeling, transferring, and applying human behaviors to both virtual and physical robots. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Peishan Cong , Ziyi Wang , Yuexin Ma , Xiangyu Yue

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, while generating…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Xingyu Chen

Recent success with large language models has sparked a new wave of verbal human-AI interaction. While such models support users in a variety of creative tasks, they lack the embodied nature of human interaction. Dance, as a primal form of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Alexander Okupnik , Johannes Schneider , Kyriakos Flouris

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

Sound · Computer Science 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

This paper focuses on enhancing human-agent communication by integrating spatial context into virtual agents' non-verbal behaviors, specifically gestures. Recent advances in co-speech gesture generation have primarily utilized data-driven…

Human-Computer Interaction · Computer Science 2024-08-09 Anna Deichler , Simon Alexanderson , Jonas Beskow

Despite progress in video-to-audio generation, the field focuses predominantly on mono output, lacking spatial immersion. Existing binaural approaches remain constrained by a two-stage pipeline that first generates mono audio and then…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Mengchen Zhang , Qi Chen , Tong Wu , Zihan Liu , Dahua Lin

While 3D human body modeling has received much attention in computer vision, modeling the acoustic equivalent, i.e. modeling 3D spatial audio produced by body motion and speech, has fallen short in the community. To close this gap, we…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Xudong Xu , Dejan Markovic , Jacob Sandakly , Todd Keebler , Steven Krenn , Alexander Richard

Text-driven human motion generation has recently attracted considerable attention, allowing models to generate human motions based on textual descriptions. However, current methods neglect the influence of human attributes-such as age,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Xinghan Wang , Kun Xu , Fei Li , Cao Sheng , Jiazhong Yu , Yadong Mu

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized Variational…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Sohan Anisetty , James Hays

Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches assume identity-neutral motion and generate movements using…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Wenqi Jia , Zekun Li , Abhay Mittal , Chengcheng Tang , Chuan Guo , Lezi Wang , James Matthew Rehg , Lingling Tao , Size An

Humans inhabit a world defined by interactions -- with other humans, objects, and environments. These interactive movements not only convey our relationships with our surroundings but also demonstrate how we perceive and communicate with…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Kewei Sui , Anindita Ghosh , Inwoo Hwang , Bing Zhou , Jian Wang , Chuan Guo

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

Graphics · Computer Science 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang

Binaural stereo audio is recorded by imitating the way the human ear receives sound, which provides people with an immersive listening experience. Existing approaches leverage autoencoders and directly exploit visual spatial information to…

Sound · Computer Science 2023-11-15 Zhaojian Li , Bin Zhao , Yuan Yuan

Human motion synthesis is an important task in computer graphics and computer vision. While focusing on various conditioning signals such as text, action class, or audio to guide the generation process, most existing methods utilize…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Kebing Xue , Hyewon Seo