English
Related papers

Related papers: DGFNet: End-to-End Audio-Visual Source Separation …

200 papers

Semantic segmentation relying solely on RGB data often struggles in challenging conditions such as low illumination and obscured views, limiting its reliability in critical applications like autonomous driving. To address this, integrating…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Ce Zhang , Zifu Wan , Simon Stepputtis , Katia Sycara , Yaqi Xie

Multi-Focus Image Fusion seeks to improve the quality of an acquired burst of images with different focus planes. For solving the task, an activity level measurement and a fusion rule are typically established to select and fuse the most…

Computer Vision and Pattern Recognition · Computer Science 2019-08-30 Fidel Alejandro Guerrero Peña , Pedro Diamel Marrero Fernández , Tsang Ing Ren , Germano Crispim Vasconcelos , Alexandre Cunha

Multi-source data classification is a critical yet challenging task for remote sensing image interpretation. Existing methods lack adaptability to diverse land cover types when modeling frequency domain features. To this end, we propose a…

Image and Video Processing · Electrical Eng. & Systems 2025-07-08 Yikang Zhao , Feng Gao , Xuepeng Jin , Junyu Dong , Qian Du

RGB-guided depth completion aims at predicting dense depth maps from sparse depth measurements and corresponding RGB images, where how to effectively and efficiently exploit the multi-modal information is a key issue. Guided dynamic…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Yufei Wang , Yuxin Mao , Qi Liu , Yuchao Dai

We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple audio sources…

Computer Vision and Pattern Recognition · Computer Science 2021-08-27 Sagnik Majumder , Ziad Al-Halah , Kristen Grauman

Source detection in graphs has demonstrated robust efficacy in the domain of rumor source identification. Although recent solutions have enhanced performance by leveraging deep neural networks, they often require complete user data. In this…

Social and Information Networks · Computer Science 2024-05-31 Le Cheng , Peican Zhu , Keke Tang , Chao Gao , Zhen Wang

This paper presents DFR (Decompose, Fuse and Reconstruct), a novel framework that addresses the fundamental challenge of effectively utilizing multi-modal guidance in few-shot segmentation (FSS). While existing approaches primarily rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Shuai Chen , Fanman Meng , Xiwei Zhang , Haoran Wei , Chenhao Wu , Qingbo Wu , Hongliang Li

Multi-Modal Relation Extraction (MMRE) aims at identifying the relation between two entities in texts that contain visual clues. Rich visual content is valuable for the MMRE task, but existing works cannot well model finer associations…

Computation and Language · Computer Science 2023-06-21 Qian Li , Shu Guo , Cheng Ji , Xutan Peng , Shiyao Cui , Jianxin Li

Extensive work has demonstrated the effectiveness of Vision Transformers. The plain Vision Transformer tends to obtain multi-scale features by selecting fixed layers, or the last layer of features aiming to achieve higher performance in…

Computer Vision and Pattern Recognition · Computer Science 2023-05-10 Fangjian Lin , Yizhe Ma , Shengwei Tian

This paper describes the deepfake audio detection system submitted to the Audio Deep Synthesis Detection (ADD) Challenge Track 3.2 and gives an analysis of score fusion. The proposed system is a score-level fusion of several light…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-14 Yuxiang Zhang , Jingze Lu , Xingming Wang , Zhuo Li , Runqiu Xiao , Wenchao Wang , Ming Li , Pengyuan Zhang

Multimodal emotion recognition has recently gained much attention since it can leverage diverse and complementary relationships over multiple modalities (e.g., audio, visual, biosignals, etc.), and can provide some robustness to noisy…

Many previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse musical video data, voice activity detection is a necessary…

Sound · Computer Science 2021-06-23 Yuanbo Hou , Zhesong Yu , Xia Liang , Xingjian Du , Bilei Zhu , Zejun Ma , Dick Botteldooren

In this work, we propose a training algorithm for an audio-visual automatic speech recognition (AV-ASR) system using deep recurrent neural network (RNN).First, we train a deep RNN acoustic model with a Connectionist Temporal Classification…

Computer Vision and Pattern Recognition · Computer Science 2016-11-10 Abhinav Thanda , Shankar M Venkatesan

Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Conventional contact-based techniques, like polygraph devices, rely on physiological signals to determine…

Training vision-language models on cognitively-plausible amounts of data requires rethinking how models integrate multimodal information. Within the constraints of the Vision track for the BabyLM Challenge 2025, we propose a lightweight…

Artificial Intelligence · Computer Science 2025-10-10 Bianca-Mihaela Ganescu , Suchir Salhan , Andrew Caines , Paula Buttery

In computer vision and image processing tasks, image fusion has evolved into an attractive research field. However, recent existing image fusion methods are mostly built on pixel-level operations, which may produce unacceptable artifacts…

Computer Vision and Pattern Recognition · Computer Science 2019-05-29 Lihua Jian , Xiaomin Yang , Zheng Liu , Gwanggil Jeon , Mingliang Gao , David Chisholm

End-to-end acoustic speech recognition has quickly gained widespread popularity and shows promising results in many studies. Specifically the joint transformer/CTC model provides very good performance in many tasks. However, under noisy and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-20 Wentao Yu , Steffen Zeiler , Dorothea Kolossa

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-03 Hiroshi Sato , Tsubasa Ochiai , Keisuke Kinoshita , Marc Delcroix , Tomohiro Nakatani , Shoko Araki

Fake news detection has received increasing attention from researchers in recent years, especially multi-modal fake news detection containing both text and images. However, many previous works have fed two modal features, text and image,…

Multimedia · Computer Science 2024-07-02 Hongzhen Lv , Wenzhong Yang , Fuyuan Wei , Jiaren Peng , Haokun Geng

We address the problem of multi-modal object tracking in video and explore various options of fusing the complementary information conveyed by the visible (RGB) and thermal infrared (TIR) modalities including pixel-level, feature-level and…

Computer Vision and Pattern Recognition · Computer Science 2022-01-24 Zhangyong Tang , Tianyang Xu , Hui Li , Xiao-Jun Wu , Xuefeng Zhu , Josef Kittler