English
Related papers

Related papers: DAVE: A Deep Audio-Visual Embedding for Dynamic Sa…

200 papers

We introduce Perception Encoder Audiovisual, PE-AV, a new family of encoders for audio and video understanding trained with scaled contrastive learning. Built on PE, PE-AV makes several key contributions to extend representations to audio,…

The goal of Audio-Visual Segmentation (AVS) is to localize and segment the sounding source objects from video frames. Research on AVS suffers from data scarcity due to the high cost of fine-grained manual annotations. Recent works attempt…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Kyungbok Lee , You Zhang , Zhiyao Duan

Sound-guided object segmentation has drawn considerable attention for its potential to enhance multimodal perception. Previous methods primarily focus on developing advanced architectures to facilitate effective audio-visual interactions,…

Sound · Computer Science 2025-03-18 Chen Liu , Liying Yang , Peike Li , Dadong Wang , Lincheng Li , Xin Yu

Deep convolutional neural networks are known to specialize in distilling compact and robust prior from a large amount of data. We are interested in applying deep networks in the absence of training dataset. In this paper, we introduce deep…

Sound · Computer Science 2019-12-24 Yapeng Tian , Chenliang Xu , Dingzeyu Li

The existing state-of-the-art method for audio-visual conditioned video prediction uses the latent codes of the audio-visual frames from a multimodal stochastic network and a frame encoder to predict the next visual frame. However, a direct…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Yating Xu , Conghui Hu , Gim Hee Lee

There has been profound progress in visual saliency thanks to the deep learning architectures, however, there still exist three major challenges that hinder the detection performance for scenes with complex compositions, multiple salient…

Computer Vision and Pattern Recognition · Computer Science 2017-08-16 Jing Zhang , Yuchao Dai , Fatih Porikli , Mingyi He

We propose DAVIS, a Diffusion-based Audio-VIsual Separation framework that solves the audio-visual sound source separation task through generative learning. Existing methods typically frame sound separation as a mask-based regression…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Chao Huang , Susan Liang , Yapeng Tian , Anurag Kumar , Chenliang Xu

Building on the success of diffusion models in image generation and editing, video editing has recently gained substantial attention. However, maintaining temporal consistency and motion alignment still remains challenging. To address these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Yi Huang , Wei Xiong , He Zhang , Chaoqi Chen , Jianzhuang Liu , Mingfu Yan , Shifeng Chen

This paper introduces the Procedural (audio) Variational autoEncoder (ProVE) framework as a general approach to learning Procedural Audio PA models of environmental sounds with an improvement to the realism of the synthesis while…

Sound · Computer Science 2023-03-07 Danzel Serrano , Mark Cartwright

Image captioning has been recently gaining a lot of attention thanks to the impressive achievements shown by deep captioning architectures, which combine Convolutional Neural Networks to extract image representations, and Recurrent Neural…

Computer Vision and Pattern Recognition · Computer Science 2018-05-22 Marcella Cornia , Lorenzo Baraldi , Giuseppe Serra , Rita Cucchiara

Speech enhancement (SE) aims to improve the quality and intelligibility of speech in noisy environments. Recent studies have shown that incorporating visual cues in audio signal processing can enhance SE performance. Given that human speech…

Sound · Computer Science 2025-05-27 Meng-Ping Lin , Jen-Cheng Hou , Chia-Wei Chen , Shao-Yi Chien , Jun-Cheng Chen , Xugang Lu , Yu Tsao

We introduce ViDaS, a two-stream, fully convolutional Video, Depth-Aware Saliency network to address the problem of attention modeling ``in-the-wild", via saliency prediction in videos. Contrary to existing visual saliency approaches using…

Computer Vision and Pattern Recognition · Computer Science 2023-05-22 Ioanna Diamanti , Antigoni Tsiami , Petros Koutras , Petros Maragos

Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where events may co-occur and exhibit varying durations. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Ling Xing , Hongyu Qu , Rui Yan , Xiangbo Shu , Jinhui Tang

To detect salient objects accurately, existing methods usually design complex backbone network architectures to learn and fuse powerful features. However, the saliency inference module that performs saliency prediction from the fused…

Computer Vision and Pattern Recognition · Computer Science 2019-03-26 Zun Li , Congyan Lang , Yunpeng Chen , Junhao Liew , Jiashi Feng

Speech enhancement systems are typically trained using pairs of clean and noisy speech. In audio-visual speech enhancement (AVSE), there is not as much ground-truth clean data available; most audio-visual datasets are collected in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-05 Ju-Chieh Chou , Chung-Ming Chien , Karen Livescu

This paper proposes a deep learning model to efficiently detect salient regions in videos. It addresses two important issues: (1) deep video saliency model training with the absence of sufficiently large and pixel-wise annotated video data,…

Computer Vision and Pattern Recognition · Computer Science 2017-12-12 Wenguan Wang , Jianbing Shen , Ling Shao

Audio-visual event (AVE) localization has attracted much attention in recent years. Most existing methods are often limited to independently encoding and classifying each video segment separated from the full video (which can be regarded as…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Yuanyuan Jiang , Jianqin Yin , Yonghao Dang

The Audio-Visual Event Localization (AVEL) task aims to temporally locate and classify video events that are both audible and visible. Most research in this field assumes a closed-set setting, which restricts these models' ability to handle…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Jinxing Zhou , Dan Guo , Ruohao Guo , Yuxin Mao , Jingjing Hu , Yiran Zhong , Xiaojun Chang , Meng Wang

We propose a novel deep neural network architecture for speech recognition that explicitly employs knowledge of the background environmental noise within a deep neural network acoustic model. A deep neural network is used to predict the…

Computation and Language · Computer Science 2016-10-03 Suyoun Kim , Bhiksha Raj , Ian Lane

Visual saliency models have enjoyed a big leap in performance in recent years, thanks to advances in deep learning and large scale annotated data. Despite enormous effort and huge breakthroughs, however, models still fall short in reaching…

Computer Vision and Pattern Recognition · Computer Science 2019-05-28 Ali Borji