English
Related papers

Related papers: How deep is your encoder: an analysis of features …

200 papers

We propose the ViNet architecture for audio-visual saliency prediction. ViNet is a fully convolutional encoder-decoder architecture. The encoder uses visual features from a network trained for action recognition, and the decoder infers a…

Computer Vision and Pattern Recognition · Computer Science 2021-08-10 Samyak Jain , Pradeep Yarlagadda , Shreyank Jyoti , Shyamgopal Karthik , Ramanathan Subramanian , Vineet Gandhi

Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 You Qin , Kai Liu , Shengqiong Wu , Kai Wang , Shijian Deng , Yapeng Tian , Junbin Xiao , Yazhou Xing , Yinghao Ma , Bobo Li , Roger Zimmermann , Lei Cui , Furu Wei , Jiebo Luo , Hao Fei

In recent years, deep neural networks have yielded state-of-the-art performance on several tasks. Although some recent works have focused on combining deep learning with recommendation, we highlight three issues of existing models. First,…

Machine Learning · Computer Science 2018-12-20 Qibing Li , Xiaolin Zheng , Xinyue Wu

We present a deep learning based methodology for extracting the singing voice signal from a musical mixture based on the underlying linguistic content. Our model follows an encoder decoder architecture and takes as input the magnitude…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-18 Pritish Chandna , Merlijn Blaauw , Jordi Bonada , Emilia Gomez

Deepfake detection is a critical task in identifying manipulated multimedia content. In real-world scenarios, deepfake content can manifest across multiple modalities, including audio and video. To address this challenge, we present…

Artificial Intelligence · Computer Science 2025-12-04 Xin Zhang , Jiaming Chu , Jian Zhao , Yuchu Jiang , Xu Yang , Lei Jin , Chi Zhang , Xuelong Li

Monocular Depth Estimation (MDE) is a fundamental computer vision task with important applications in 3D vision. The current mainstream MDE methods employ an encoder-decoder architecture with multi-level/scale feature processing. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Huibin Bai , Shuai Li , Hanxiao Zhai , Yanbo Gao , Chong Lv , Yibo Wang , Haipeng Ping , Wei Hua , Xingyu Gao

Image Quality Assessment algorithms predict a quality score for a pristine or distorted input image, such that it correlates with human opinion. Traditional methods required a non-distorted "reference" version of the input image to compare…

Image and Video Processing · Electrical Eng. & Systems 2020-07-21 Subhayan Mukherjee , Giuseppe Valenzise , Irene Cheng

Variational autoencoders (VAEs) have been used extensively to discover low-dimensional latent factors governing neural activity and animal behavior. However, without careful model selection, the uncovered latent factors may reflect noise in…

Machine Learning · Computer Science 2023-12-13 Julia Huiming Wang , Dexter Tsin , Tatiana Engel

Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse questions about audio-visual scenes. However, existing methods lack sufficient flexibility and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Jiayu Zhang , Shuo Ye , Qilang Ye , Xun Lin , Zihan Song , Zitong Yu

Inspired by the success of deep neural networks (DNNs) in speech processing, this paper presents Deep Vocoder, a direct end-to-end low bit rate speech compression method with deep autoencoder (DAE). In Deep Vocoder, DAE is used for…

Multimedia · Computer Science 2019-05-15 Gang Min , Changqing Zhang , Xiongwei Zhang , Wei Tan

The expressive nature of the voice provides a powerful medium for communicating sonic ideas, motivating recent research on methods for query by vocalisation. Meanwhile, deep learning methods have demonstrated state-of-the-art results for…

Multimedia · Computer Science 2018-02-15 Adib Mehrabi , Keunwoo Choi , Simon Dixon , Mark Sandler

Efficiently representing audio signals in a compressed latent space is critical for latent generative modelling. However, existing autoencoders often force a choice between continuous embeddings and discrete tokens. Furthermore, achieving…

Sound · Computer Science 2025-09-15 Marco Pasini , Stefan Lattner , George Fazekas

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Qiang Huang , Thomas Hain

Depth estimation and 3D reconstruction have been extensively studied as core topics in computer vision. Starting from rigid objects with relatively simple geometric shapes, such as vehicles, the research has expanded to address general…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Muhammad Aamir , Naoya Muramatsu , Sangyun Shin , Matthew Wijers , Jia-Xing Zhong , Xinyu Hou , Amir Patel , Andrew Loveridge , Andrew Markham

Audio-visual segmentation (AVS) plays a critical role in multimodal machine learning by effectively integrating audio and visual cues to precisely segment objects or regions within visual scenes. Recent AVS methods have demonstrated…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Yunzhe Shen , Kai Peng , Leiye Liu , Wei Ji , Jingjing Li , Miao Zhang , Yongri Piao , Huchuan Lu

Dialog systems need to understand dynamic visual scenes in order to have conversations with users about the objects and events around them. Scene-aware dialog systems for real-world applications could be developed by integrating…

Transformer-based end-to-end speech recognition has achieved great success. However, the large footprint and computational overhead make it difficult to deploy these models in some real-world applications. Model compression techniques can…

Computation and Language · Computer Science 2023-03-15 Yifan Peng , Jaesong Lee , Shinji Watanabe

Deep learning-based quality assessments have significantly enhanced perceptual multimedia quality assessment, however it is still in the early stages for 3D visual data such as 3D point clouds (PCs). Due to the high volume of 3D-PCs, such…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Oussama Messai , Abdelouahid Bentamou , Abbass Zein-Eddine , Yann Gavet

Visual-to-auditory sensory substitution devices can assist the blind in sensing the visual environment by translating the visual information into a sound pattern. To improve the translation quality, the task performances of the blind are…

Computer Vision and Pattern Recognition · Computer Science 2019-04-22 Di Hu , Dong Wang , Xuelong Li , Feiping Nie , Qi Wang

Automatic speaker verification (ASV) is the process to recognize persons using voice as biometric. The ASV systems show considerable recognition performance with sufficient amount of speech from matched condition. One of the crucial…

Multimedia · Computer Science 2018-12-04 Arnab Poddar , Md Sahidullah , Goutam Saha
‹ Prev 1 8 9 10 Next ›