English
Related papers

Related papers: FolAI: Synchronized Foley Sound Generation with Se…

200 papers

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-scale or complex…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Zijun Wang , Panwen Hu , Jing Wang , Terry Jingchen Zhang , Yuhao Cheng , Long Chen , Yiqiang Yan , Zutao Jiang , Hanhui Li , Xiaodan Liang

Speech synthesis has recently seen significant improvements in fidelity, driven by the advent of neural vocoders and neural prosody generators. However, these systems lack intuitive user controls over prosody, making them unable to rectify…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-13 Max Morrison , Zeyu Jin , Justin Salamon , Nicholas J. Bryan , Gautham J. Mysore

A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instantaneous velocity,…

Sound · Computer Science 2025-09-09 Xiaoran Yang , Jianxuan Yang , Xinyue Guo , Haoyu Wang , Ningning Pan , Gongping Huang

Video-to-music generation presents significant potential in video production, requiring the generated music to be both semantically and rhythmically aligned with the video. Achieving this alignment demands advanced music generation…

Sound · Computer Science 2024-12-10 Sifei Li , Binxin Yang , Chunji Yin , Chong Sun , Yuxin Zhang , Weiming Dong , Chen Li

Audio-driven talking face generation has received growing interest, particularly for applications requiring expressive and natural human-avatar interaction. However, most existing emotion-aware methods rely on a single modality (either…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Phyo Thet Yee , Dimitrios Kollias , Sudeepta Mishra , Abhinav Dhall

Spatial audio is crucial for immersive 360-degree video experiences, yet most 360-degree videos lack it due to the difficulty of capturing spatial audio during recording. Automatically generating spatial audio such as first-order ambisonics…

Sound · Computer Science 2026-04-13 Ziyu Luo , Lin Chen , Qiang Qu , Xiaoming Chen , Yiran Shen

Spatial audio is crucial for immersive 360-degree video experiences, yet most 360-degree videos lack it due to the difficulty of capturing spatial audio during recording. Automatically generating spatial audio such as first-order ambisonics…

Sound · Computer Science 2026-05-05 Ziyu Luo , Lin Chen , Qiang Qu , Xiaoming Chen , Yiran Shen

The field has made significant progress in synthesizing realistic human motion driven by various modalities. Yet, the need for different methods to animate various body parts according to different control signals limits the scalability of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Zixiang Zhou , Yu Wan , Baoyuan Wang

Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. Controlling stereo…

Sound · Computer Science 2025-02-26 Peiwen Sun , Sitong Cheng , Xiangtai Li , Zhen Ye , Huadai Liu , Honggang Zhang , Wei Xue , Yike Guo

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-matching generative…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Peiyin Chen , Zhuowei Yang , Hui Feng , Sheng Jiang , Rui Yan

Generating realistic, context-aware two-person motion conditioned on diverse modalities remains a fundamental challenge for graphics, animation and embodied AI systems. Real-world applications such as VR/AR companions, social robotics and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Prerit Gupta , Shourya Verma , Ananth Grama , Aniket Bera

Hand-Object Interaction (HOI) remains a core challenge in digital human video synthesis, where models must generate physically plausible contact and preserve object identity across frames. Although recent HOI reenactment approaches have…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Xuan Huang , Mochu Xiang , Zhelun Shen , Jinbo Wu , Chenming Wu , Chen Zhao , Kaisiyuan Wang , Hang Zhou , Shanshan Liu , Haocheng Feng , Wei He , Jingdong Wang

Audio is an essential part of our life, but creating it often requires expertise and is time-consuming. Research communities have made great progress over the past year advancing the performance of large scale audio generative models for a…

Recent advancements in video generation models have significantly improved their ability to follow text prompts. However, the customization of dynamic visual effects, defined as temporally evolving and appearance-driven visual phenomena…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Rui Zhao , Mike Zheng Shou

Video-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency, and visual-audio temporal synchrony. We propose Frieren, a…

Sound · Computer Science 2025-01-07 Yongqi Wang , Wenxiang Guo , Rongjie Huang , Jiawei Huang , Zehan Wang , Fuming You , Ruiqi Li , Zhou Zhao

Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjusting the style while…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-01 Hanzhao Li , Yuke Li , Xinsheng Wang , Jingbin Hu , Qicong Xie , Shan Yang , Lei Xie

Music generation in the audio domain using artificial intelligence (AI) has witnessed steady progress in recent years. However for some instruments, particularly the guitar, controllable instrument synthesis remains limited in expressivity.…

Sound · Computer Science 2025-10-28 Jackson Loth , Pedro Sarmento , Mark Sandler , Mathieu Barthet

Vocals harmonizers are powerful tools to help solo vocalists enrich their melodies with harmonically supportive voices. These tools exist in various forms, from commercially available pedals and software to custom-built systems, each…

Human-Computer Interaction · Computer Science 2025-06-24 Lancelot Blanchard , Cameron Holt , Joseph A. Paradiso

The field of text-to-audio generation has seen significant advancements, and yet the ability to finely control the acoustic characteristics of generated audio remains under-explored. In this paper, we introduce a novel yet simple approach…

Sound · Computer Science 2024-12-16 Sonal Kumar , Prem Seetharaman , Justin Salamon , Dinesh Manocha , Oriol Nieto

Insufficient recordings and the scarcity of anomalies present significant challenges in developing and validating robust anomaly detection systems for machine sounds. To address these limitations, we propose a novel approach for generating…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-30 Harsh Purohit , Tomoya Nishida , Kota Dohi , Takashi Endo , Yohei Kawaguchi