English
Related papers

Related papers: SeamlessEdit: Background Noise Aware Zero-Shot Spe…

200 papers

Audio editing plays a central role in VR/AR immersion, virtual conferencing, sound design, and other interactive media. However, recent generative audio editing models depend on template-like instruction formats and are restricted to…

Sound · Computer Science 2025-09-29 Zitong Lan , Yiduo Hao , Mingmin Zhao

In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Yan-Bo Lin , Kevin Lin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Chung-Ching Lin , Xiaofei Wang , Gedas Bertasius , Lijuan Wang

In this paper, we introduce a novel task called language-guided joint audio-visual editing. Given an audio and image pair of a sounding event, this task aims at generating new audio-visual content by editing the given sounding event…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Susan Liang , Chao Huang , Yapeng Tian , Anurag Kumar , Chenliang Xu

Internal features from large-scale pre-trained diffusion models have recently been established as powerful semantic descriptors for a wide range of downstream tasks. Works that use these features generally need to add noise to images before…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Nick Stracke , Stefan Andreas Baumann , Kolja Bauer , Frank Fundel , Björn Ommer

Voice Conversion research in recent times has increasingly focused on improving the zero-shot capabilities of existing methods. Despite remarkable advancements, current architectures still tend to struggle in zero-shot cross-lingual…

Sound · Computer Science 2025-05-26 Advait Joglekar , Divyanshu Singh , Rooshil Rohit Bhatia , S. Umesh

Speech clarity and spatial audio immersion are the two most critical factors in enhancing remote conferencing experiences. Existing methods are often limited: either due to the lack of spatial information when using only one microphone, or…

Sound · Computer Science 2025-07-14 Cheng Chi , Xiaoyu Li , Yuxuan Ke , Qunping Ni , Yao Ge , Xiaodong Li , Chengshi Zheng

Generative Universal Speech Enhancement (USE) methods aim to leverage generative models to improve speech quality under various types of distortions. However, existing generative speech enhancement methods often suffer from semantic…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-07 Xingchen Li , Hanke Xie , Ziqian Wang , Zihan Zhang , Longshuai Xiao , Shuai Wang , Lei Xie

The quantification of audio aesthetics remains a complex challenge in audio processing, primarily due to its subjective nature, which is influenced by human perception and cultural context. Traditional methods often depend on human…

Speech enhancement (SE) is usually required as a front end to improve the speech quality in noisy environments, while the enhanced speech might not be optimal for automatic speech recognition (ASR) systems due to speech distortion. On the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-27 Qiu-Shi Zhu , Jie Zhang , Zi-Qiang Zhang , Li-Rong Dai

A Spoken dialogue system for an unseen language is referred to as Zero resource speech. It is especially beneficial for developing applications for languages that have low digital resources. Zero resource speech synthesis is the task of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-11 Karthik Pandia D S , Anusha Prakash , Mano Ranjith Kumar , Hema A Murthy

We propose a method for adding sound-guided visual effects to specific regions of videos with a zero-shot setting. Animating the appearance of the visual effect is challenging because each frame of the edited video should have visual…

Computer Vision and Pattern Recognition · Computer Science 2023-04-17 Seung Hyun Lee , Sieun Kim , Innfarn Yoo , Feng Yang , Donghyeon Cho , Youngseo Kim , Huiwen Chang , Jinkyu Kim , Sangpil Kim

Recent advances in text-to-image models have increased the exposure of powerful image editing techniques as a tool, raising concerns about their potential for malicious use. An emerging line of research to address such threats focuses on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Jinsu Kim , Yunhun Nam , Minseon Kim , Sangpil Kim , Jongheon Jeong

Taking inspiration from recent developments in visual generative tasks using diffusion models, we propose a method for end-to-end speech-driven video editing using a denoising diffusion model. Given a video of a talking person, and a…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Dan Bigioi , Shubhajit Basak , Michał Stypułkowski , Maciej Zięba , Hugh Jordan , Rachel McDonnell , Peter Corcoran

Spatial audio understanding is essential for accurately perceiving and interpreting acoustic environments. However, existing audio-language models exhibit limitations in processing spatial audio and perceiving spatial acoustic scenes. To…

Sound · Computer Science 2025-09-19 Jinbo Hu , Yin Cao , Ming Wu , Zhenbo Luo , Jun Yang

Text-to-image diffusion models have achieved remarkable success in generating high-quality and diverse images. Building on these advancements, diffusion models have also demonstrated exceptional performance in text-guided image editing. A…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Mingyu Kang , Yong Suk Choi

Anomalous Sound Detection (ASD) aims at identifying anomalous sounds from machines and has gained extensive research interests from both academia and industry. However, the uncertainty of anomaly location and much redundant information such…

Sound · Computer Science 2025-08-22 Guirui Zhong , Qing Wang , Jun Du , Lei Wang , Mingqi Cai , Xin Fang

Music editing has emerged as an important and practical area of artificial intelligence, with applications ranging from video game and film music production to personalizing existing tracks according to user preferences. However, existing…

Sound · Computer Science 2025-11-19 Ali Boudaghi , Hadi Zare

Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms directly from the text. Single-stage text-to-speech often faces…

Sound · Computer Science 2022-07-14 Zhengxi Liu , Qiao Tian , Chenxu Hu , Xudong Liu , Menglin Wu , Yuping Wang , Hang Zhao , Yuxuan Wang

In this paper, we explore an improved framework to train a monoaural neural enhancement model for robust speech recognition. The designed training framework extends the existing mixture invariant training criterion to exploit both unpaired…

Sound · Computer Science 2022-09-21 Jisi Zhang , Catalin Zorila , Rama Doddipatla , Jon Barker

Given a pair of source and reference speech recordings, speech-to-speech (S2S) emotion style transfer involves the generation of an output speech that mimics the emotion characteristics of the reference while preserving the content and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Soumya Dutta , Avni Jain , Sriram Ganapathy