中文
相关论文

相关论文: PAVAS: Physics-Aware Video-to-Audio Synthesis

200 篇论文

Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data. However, audio recordings with high-quality text annotations can be difficult to acquire. In this work, we approach text-to-audio synthesis using…

While persona-driven large language models (LLMs) and prompt-based text-to-speech (TTS) systems have advanced significantly, a usability gap arises when users attempt to generate voices matching their desired personas from implicit…

音频与语音处理 · 电气工程与系统科学 2025-09-22 Yejin Lee , Jaehoon Kang , Kyuhong Shim

Generating speech-consistent body and gesture movements is a long-standing problem in virtual avatar creation. Previous studies often synthesize pose movement in a holistic manner, where poses of all joints are generated simultaneously.…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Xian Liu , Qianyi Wu , Hang Zhou , Yinghao Xu , Rui Qian , Xinyi Lin , Xiaowei Zhou , Wayne Wu , Bo Dai , Bolei Zhou

Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges: (1) the scarcity…

声音 · 计算机科学 2026-04-30 Yusheng Dai , Zehua Chen , Yuxuan Jiang , Baolong Gao , Qiuhong Ke , Jianfei Cai , Jun Zhu

Speech is a means of communication which relies on both audio and visual information. The absence of one modality can often lead to confusion or misinterpretation of information. In this paper we present an end-to-end temporal model capable…

音频与语音处理 · 电气工程与系统科学 2019-06-17 Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

Recent advances in generative AI offer promising solutions for synthetic data generation but often rely on large datasets for effective training. To address this limitation, we propose a novel generative model that learns from limited data…

机器学习 · 统计学 2025-05-27 Michail Spitieris , Massimiliano Ruocco , Abdulmajid Murad , Alessandro Nocente

In recent years, there has been rapid development in 3D generation models, opening up new possibilities for applications such as simulating the dynamic movements of 3D objects and customizing their behaviors. However, current 3D generative…

计算机视觉与模式识别 · 计算机科学 2024-06-12 Fangfu Liu , Hanyang Wang , Shunyu Yao , Shengjun Zhang , Jie Zhou , Yueqi Duan

Video-to-Audio generation has made remarkable strides in automatically synthesizing sound for video. However, existing evaluation metrics, which focus on semantic and temporal alignment, overlook a critical failure mode: models often…

声音 · 计算机科学 2025-12-29 Liyang Chen , Hongkai Chen , Yujun Cai , Sifan Li , Qingwen Ye , Yiwei Wang

With recent advances of AIGC, video generation have gained a surge of research interest in both academia and industry (e.g., Sora). However, it remains a challenge to produce temporally aligned audio to synchronize the generated video,…

音频与语音处理 · 电气工程与系统科学 2024-09-24 Yuchen Hu , Yu Gu , Chenxing Li , Rilin Chen , Dong Yu

Modeling and rendering photorealistic avatars is of crucial importance in many applications. Existing methods that build a 3D avatar from visual observations, however, struggle to reconstruct clothed humans. We introduce PhysAvatar, a novel…

Spatial audio is fundamental to immersive virtual experiences, yet synthesizing high-fidelity binaural audio from sparse observations remains a significant challenge. Existing methods typically rely on implicit neural representations…

声音 · 计算机科学 2026-04-13 Chunhao Bi , Houqiang Zhong , Zhixin Xu , Li Song , Zhengxue Cheng

We present a novel audio-driven facial animation approach that can generate realistic lip-synchronized 3D facial animations from the input audio. Our approach learns viseme dynamics from speech videos, produces animator-friendly viseme…

图形学 · 计算机科学 2023-01-18 Linchao Bao , Haoxian Zhang , Yue Qian , Tangli Xue , Changhai Chen , Xuefei Zhe , Di Kang

As embodied agents become central to VR, telepresence, and digital human applications, their motion must go beyond speech-aligned gestures: agents should turn toward users, respond to their movement, and maintain natural gaze. Current…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Evonne Ng , Siwei Zhang , Zhang Chen , Michael Zollhoefer , Alexander Richard

Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal…

声音 · 计算机科学 2024-09-16 Zhiqi Huang , Dan Luo , Jun Wang , Huan Liao , Zhiheng Li , Zhiyong Wu

In this study, we introduce AV-PedAware, a self-supervised audio-visual fusion system designed to improve dynamic pedestrian awareness for robotics applications. Pedestrian awareness is a critical requirement in many robotics applications.…

机器人学 · 计算机科学 2025-04-07 Yizhuo Yang , Shenghai Yuan , Muqing Cao , Jianfei Yang , Lihua Xie

We present PhysGen, a novel image-to-video generation method that converts a single image and an input condition (e.g., force and torque applied to an object in the image) to produce a realistic, physically plausible, and temporally…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Shaowei Liu , Zhongzheng Ren , Saurabh Gupta , Shenlong Wang

While expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we here address the control of perceptual voice qualities…

音频与语音处理 · 电气工程与系统科学 2025-01-16 Frederik Rautenberg , Michael Kuhlmann , Fritz Seebauer , Jana Wiechmann , Petra Wagner , Reinhold Haeb-Umbach

Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speech, while allowing flexible control of speaker identity, all…

音频与语音处理 · 电气工程与系统科学 2022-02-21 Disong Wang , Shan Yang , Dan Su , Xunying Liu , Dong Yu , Helen Meng

We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-to-audio models achieve strong semantic…

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

声音 · 计算机科学 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li