English
Related papers

Related papers: How deep is your encoder: an analysis of features …

200 papers

Metric depth prediction from monocular videos suffers from bad generalization between datasets and requires supervised depth data for scale-correct training. Self-supervised training using multi-view reconstruction can benefit from large…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Xiaohu Liu , Sascha Hornauer , Fabien Moutarde , Jialiang Lu

Neural audio codecs (NACs) achieve low-bitrate compression by learning compact audio representations, which can also serve as features for perceptual quality evaluation. We introduce DACe, an enhanced, higher-fidelity version of the…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Arijit Biswas , Lars Villemoes

Self-supervised learning has become a central strategy for representation learning, but the majority of architectures used for encoding data have only been validated on regularly-sampled inputs such as images, audios. and videos. In many…

Machine Learning · Statistics 2025-10-24 Yunyi Shen , Alexander Gagliano

Video-based dialog task is a challenging multimodal learning task that has received increasing attention over the past few years with state-of-the-art obtaining new performance records. This progress is largely powered by the adaptation of…

Computer Vision and Pattern Recognition · Computer Science 2022-10-27 Huda Alamri , Anthony Bilic , Michael Hu , Apoorva Beedu , Irfan Essa

Video saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip. Existing VSD-related works have mainly relied on the visual system but paid less attention to the audio aspect, while,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Chenglizhao Chen , Mengke Song , Wenfeng Song , Li Guo , Muwei Jian

Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information. However, a thorough understanding of videos significantly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yudong Yang , Jimin Zhuang , Guangzhi Sun , Changli Tang , Yixuan Li , Peihan Li , Yifan Jiang , Wei Li , Zejun Ma , Chao Zhang

We introduce a new convolutional AutoEncoder architecture for user modelling and recommendation tasks with several improvements over the state of the art. Firstly, our model has the flexibility to learn a set of associations and…

Machine Learning · Computer Science 2025-09-10 Antoine Ledent , Petr Kasalický , Rodrigo Alves , Hady W. Lauw

This paper introduces a novel audio-to-image encoding framework that integrates multiple dimensions of voice characteristics into a single RGB image for speaker recognition. In this method, the green channel encodes raw audio data, the red…

Sound · Computer Science 2025-03-11 Youness Atif

Accurate volume estimation of objects from visual data is a long-standing challenge in computer vision with significant applications in robotics, logistics, and smart health. Existing methods often rely on complex 3D reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Gautham Vinod , Bruce Coburn , Siddeshwar Raghavan , Fengqing Zhu

Neural network based approaches to speech enhancement have shown to be particularly powerful, being able to leverage a data-driven approach to result in a significant performance gain versus other approaches. Such approaches are reliant on…

Sound · Computer Science 2023-12-15 George Close , William Ravenscroft , Thomas Hain , Stefan Goetze

In an earlier study, we gathered perceptual evaluations of the audio, video, and audiovisual quality for 360 audiovisual content. This paper investigates perceived audiovisual quality prediction based on objective quality metrics and…

Multimedia · Computer Science 2021-12-24 Randy Frans Fela , Nick Zacharov , Søren Forchhammer

Audio-visual deepfakes have reached a level of realism that makes perceptual detection unreliable, threatening media integrity and biometric security. While multimodal detection has shown promise, most approaches are binary classification…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Wasim Ahmad , Wei Zhang , Xuerui Mao

There has been significant research effort developing neural-network-based predictors of SQ in recent years. While a primary objective has been to develop non-intrusive, i.e.~reference-free, metrics to assess the performance of SE systems,…

Sound · Computer Science 2025-08-05 George Close , Kris Hong , Thomas Hain , Stefan Goetze

With the recent advancements in Artificial Intelligence (AI), Intelligent Virtual Assistants (IVA) such as Alexa, Google Home, etc., have become a ubiquitous part of many homes. Currently, such IVAs are mostly audio-based, but going…

Multimedia · Computer Science 2019-12-27 Shachi H Kumar , Eda Okur , Saurav Sahay , Jonathan Huang , Lama Nachman

Recently, direct modeling of raw waveforms using deep neural networks has been widely studied for a number of tasks in audio domains. In speaker verification, however, utilization of raw waveforms is in its preliminary phase, requiring…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-18 Jee-weon Jung , Hee-Soo Heo , Ju-ho Kim , Hye-jin Shim , Ha-Jin Yu

In a voice-controlled smart-home, a controller must respond not only to user's requests but also according to the interaction context. This paper describes Arcades, a system which uses deep reinforcement learning to extract context from a…

Machine Learning · Computer Science 2018-07-19 Alexis Brenon , François Portet , Michel Vacher

This paper proposes a novel framework for unsupervised audio source separation using a deep autoencoder. The characteristics of unknown source signals mixed in the mixed input is automatically by properly configured autoencoders implemented…

Sound · Computer Science 2014-12-24 Giljin Jang , Han-Gyu Kim , Yung-Hwan Oh

Over the recent years, various deep learning-based embedding methods have been proposed and have shown impressive performance in speaker verification. However, as in most of the classical embedding techniques, the deep learning-based…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-10 Woo Hyun Kang , Sung Hwan Mun , Min Hyun Han , Nam Soo Kim

We present a framework to model the perceived quality of audio signals by combining convolutional architectures, with ideas from classical signal processing, and describe an approach to enhancing perceived acoustical quality. We demonstrate…

Sound · Computer Science 2019-12-13 Prateek Verma , Jonathan Berger

Audio descriptions (ADs) narrate important visual details in movies, enabling Blind and Low Vision (BLV) users to understand narratives and appreciate visual details. Existing works in automatic AD generation mostly focus on few-second…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Divy Kala , Eshika Khandelwal , Makarand Tapaswi