English
Related papers

Related papers: Sep-Stereo: Visually Guided Stereophonic Audio Gen…

200 papers

The ability to envisage the visual of a talking face based just on hearing a voice is a unique human capability. There have been a number of works that have solved for this ability recently. We differ from these approaches by enabling a…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Ravindra Yadav , Ashish Sardana , Vinay P Namboodiri , Rajesh M Hegde

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main…

Sound · Computer Science 2020-01-03 Rongzhi Gu , Yuexian Zou

We propose a step-by-step video-to-audio (V2A) generation method for finer controllability over the generation process and more realistic audio synthesis. Inspired by traditional Foley workflows, our approach aims to comprehensively capture…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Akio Hayakawa , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

The framework of visually-guided sound source separation generally consists of three parts: visual feature extraction, multimodal feature fusion, and sound signal processing. An ongoing trend in this field has been to tailor involved visual…

Sound · Computer Science 2023-06-21 Zengjie Song , Zhaoxiang Zhang

Speech signals are inherently complex as they encompass both global acoustic characteristics and local semantic information. However, in the task of target speech extraction, certain elements of global and local semantic information in the…

Sound · Computer Science 2024-08-27 Zhaoxi Mu , Xinyu Yang , Sining Sun , Qing Yang

We present SingSong, a system that generates instrumental music to accompany input vocals, potentially offering musicians and non-musicians alike an intuitive new way to create music featuring their own voice. To accomplish this, we build…

This study introduces a novel training paradigm, audio difference learning, for improving audio captioning. The fundamental concept of the proposed learning method is to create a feature representation space that preserves the relationship…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Tatsuya Komatsu , Yusuke Fujita , Kazuya Takeda , Tomoki Toda

Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing methods mostly entangle memory modeling with video generation, leading to inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Yanjun Guo , Zhengqiang Zhang , Pengfei Wang , Xinyue Liang , Zhiyuan Ma , Lei Zhang

In this paper, we propose a novel diffusion-based approach to generate stereo images given a text prompt. Since stereo image datasets with large baselines are scarce, training a diffusion model from scratch is not feasible. Therefore, we…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Aakash Garg , Libing Zeng , Andrii Tsarov , Nima Khademi Kalantari

Recently, audio generation tasks have attracted considerable research interests. Precise temporal controllability is essential to integrate audio generation with real applications. In this work, we propose a temporal controlled audio…

Sound · Computer Science 2024-07-18 Zeyu Xie , Xuenan Xu , Zhizheng Wu , Mengyue Wu

Speech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different…

Signal Processing · Electrical Eng. & Systems 2021-02-09 Nicolas Furnon , Romain Serizel , Irina Illina , Slim Essid

Sight and hearing are two senses that play a vital role in human communication and scene understanding. To mimic human perception ability, audio-visual learning, aimed at developing computational approaches to learn from both audio and…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Yake Wei , Di Hu , Yapeng Tian , Xuelong Li

Generating audio from a video's visual context has multiple practical applications in improving how we interact with audio-visual media - for example, enhancing CCTV footage analysis, restoring historical videos (e.g., silent movies), and…

Sound · Computer Science 2024-04-30 Hugo Garrido-Lestache Belinchon , Helina Mulugeta , Adam Haile

Neural upmixing, the task of generating immersive music with an increased number of channels from fewer input channels, has been an active research area, with mono-to-stereo and stereo-to-surround upmixing treated as separate problems. In…

Sound · Computer Science 2024-05-24 Yongyi Zang , Yifan Wang , Minglun Lee

Supervised deep learning approaches to underdetermined audio source separation achieve state-of-the-art performance but require a dataset of mixtures along with their corresponding isolated source signals. Such datasets can be extremely…

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interplay between audio and…

Multimedia · Computer Science 2024-10-01 Kun Su , Xiulong Liu , Eli Shlizerman

Audio editing plays a central role in VR/AR immersion, virtual conferencing, sound design, and other interactive media. However, recent generative audio editing models depend on template-like instruction formats and are restricted to…

Sound · Computer Science 2025-09-29 Zitong Lan , Yiduo Hao , Mingmin Zhao

Deep neural networks have shown excellent performance for stereo matching. Many efforts focus on the feature extraction and similarity measurement of the matching cost computation step while less attention is paid on cost aggregation which…

Computer Vision and Pattern Recognition · Computer Science 2018-01-15 Lidong Yu , Yucheng Wang , Yuwei Wu , Yunde Jia

This work pioneers the utilization of generative features in enhancing audio understanding. Unlike conventional discriminative features that directly optimize posterior and thus emphasize semantic abstraction while losing fine grained…

Sound · Computer Science 2025-09-30 Zeyu Xie , Chenxing Li , Xuenan Xu , Mengyue Wu , Wenfu Wang , Ruibo Fu , Meng Yu , Dong Yu , Yuexian Zou

Conventional approaches to sound localization and separation are based on microphone arrays in artificial systems. Inspired by the selective perception of human auditory system, we design a multi-source listening system which can separate…

Sound · Computer Science 2019-11-11 Xuecong Sun , Han Jia , Zhe Zhang , Yuzhen Yang , Zhaoyong Sun , Jun Yang
‹ Prev 1 8 9 10 Next ›