English
Related papers

Related papers: SALSA-V: Shortcut-Augmented Long-form Synchronized…

200 papers

Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data. However, audio recordings with high-quality text annotations can be difficult to acquire. In this work, we approach text-to-audio synthesis using…

Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information. In this work, we tackle the scaling challenge in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Christian Simon , Masato Ishii , Wei-Yao Wang , Koichi Saito , Akio Hayakawa , Dongseok Shim , Zhi Zhong , Shuyang Cui , Shusuke Takahashi , Takashi Shibuya , Yuki Mitsufuji

Several end-to-end deep learning approaches have been recently presented which simultaneously extract visual features from the input images and perform visual speech classification. However, research on jointly extracting audio and visual…

Computer Vision and Pattern Recognition · Computer Science 2017-09-14 Stavros Petridis , Yujiang Wang , Zuwei Li , Maja Pantic

The goal of DCASE 2023 Challenge Task 7 is to generate various sound clips for Foley sound synthesis (FSS) by "category-to-sound" approach. "Category" is expressed by a single index while corresponding "sound" covers diverse and different…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-09 Junhyeok Lee , Hyeonuk Nam , Yong-Hwa Park

Deep learning models are mostly used in an offline inference fashion. However, this strongly limits the use of these models inside audio generation setups, as most creative workflows are based on real-time digital signal processing.…

Sound · Computer Science 2022-04-15 Antoine Caillon , Philippe Esling

This paper proposes a voice conversion (VC) method using sequence-to-sequence (seq2seq or S2S) learning, which flexibly converts not only the voice characteristics but also the pitch contour and duration of input speech. The proposed…

Sound · Computer Science 2020-10-08 Hirokazu Kameoka , Kou Tanaka , Damian Kwasny , Takuhiro Kaneko , Nobukatsu Hojo

Deep learning based visual to sound generation systems essentially need to be developed particularly considering the synchronicity aspects of visual and audio features with time. In this research we introduce a novel task of guiding a class…

Machine Learning · Computer Science 2021-07-21 Sanchita Ghose , John J. Prevost

We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Sicheng Xu , Guojun Chen , Yu-Xiao Guo , Jiaolong Yang , Chong Li , Zhenyu Zang , Yizhong Zhang , Xin Tong , Baining Guo

Multi-speaker singing voice synthesis is to generate the singing voice sung by different speakers. To generalize to new speakers, previous zero-shot singing adaptation methods obtain the timbre of the target speaker with a fixed-size…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-12 Shoutong Wang , Jinglin Liu , Yi Ren , Zhen Wang , Changliang Xu , Zhou Zhao

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data, such as conflating…

Sound · Computer Science 2026-03-23 Pengjun Fang , Yingqing He , Yazhou Xing , Qifeng Chen , Ser-Nam Lim , Harry Yang

Foley sound generation, the art of creating audio for multimedia, has recently seen notable advancements through text-conditioned latent diffusion models. These systems use multimodal text-audio representation models, such as Contrastive…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-15 Tornike Karchkhadze , Hassan Salami Kavaki , Mohammad Rasool Izadi , Bryce Irvin , Mikolaj Kegler , Ari Hertz , Shuo Zhang , Marko Stamenovic

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to…

Computer Vision and Pattern Recognition · Computer Science 2022-06-15 Changan Chen , Ruohan Gao , Paul Calamia , Kristen Grauman

The ability to envisage the visual of a talking face based just on hearing a voice is a unique human capability. There have been a number of works that have solved for this ability recently. We differ from these approaches by enabling a…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Ravindra Yadav , Ashish Sardana , Vinay P Namboodiri , Rajesh M Hegde

We introduce UniVerse-1, a unified, Veo-3-like model capable of simultaneously generating coordinated audio and video. To enhance training efficiency, we bypass training from scratch and instead employ a stitching of experts (SoE)…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Duomin Wang , Wei Zuo , Aojie Li , Ling-Hao Chen , Xinyao Liao , Deyu Zhou , Zixin Yin , Xili Dai , Daxin Jiang , Gang Yu

Ultrasound video classification enables automated diagnosis and has emerged as an important research area. However, publicly available ultrasound video datasets remain scarce, hindering progress in developing effective video classification…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Tingxiu Chen , Yilei Shi , Zixuan Zheng , Bingcong Yan , Jingliang Hu , Xiao Xiang Zhu , Lichao Mou

Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model, and training that…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Vladimir Iashin , Weidi Xie , Esa Rahtu , Andrew Zisserman

Audio-visual speech separation (AVSS) aims to extract a target speech signal from a mixed signal by leveraging both auditory and visual (lip movement) cues. However, most existing AVSS methods exhibit complex architectures and rely on…

Sound · Computer Science 2025-10-15 Wendi Sang , Kai Li , Runxuan Yang , Jianqiang Huang , Xiaolin Hu

Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-08 Kaidi Wang , Yi He , Wenhao Guan , Weijie Wu , Hongwu Ding , Xiong Zhang , Di Wu , Meng Meng , Jian Luan , Lin Li , Qingyang Hong

The diffusion-based Singing Voice Conversion (SVC) methods have achieved remarkable performances, producing natural audios with high similarity to the target timbre. However, the iterative sampling process results in slow inference speed,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-04 Yiwen Lu , Zhen Ye , Wei Xue , Xu Tan , Qifeng Liu , Yike Guo

We present a lightweight latent diffusion model for vocal-conditioned musical accompaniment generation that addresses critical limitations in existing music AI systems. Our approach introduces a novel soft alignment attention mechanism that…

Sound · Computer Science 2026-01-06 Hei Shing Cheung , Boya Zhang , Jonathan H. Chan
‹ Prev 1 8 9 10 Next ›