English
Related papers

Related papers: How deep is your encoder: an analysis of features …

200 papers

Although large-scale visual foundation models (VFMs) achieve remarkable performance in semantic understanding, they still underperform in instance-aware dense prediction tasks. They exhibit different biases in representation: for instance,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yachan Guo , JoseLuis Gomez Zurita , Danna Xue , Yi Xiao , AntonioManuel Lopez Pena

Single-shot imaging with femtosecond X-ray lasers is a powerful measurement technique that can achieve both high spatial and temporal resolution. However, its accuracy has been severely limited by the difficulty of applying conventional…

Recognizing the sounding objects in scenes is a longstanding objective in embodied AI, with diverse applications in robotics and AR/VR/MR. To that end, Audio-Visual Segmentation (AVS), taking as condition an audio signal to identify the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Artem Sokolov , Swapnil Bhosale , Xiatian Zhu

What happens when we push audio-visual alignment to its absolute limits? To systematically investigate this question, we needed datasets with granular alignment quality annotations, but existing datasets treat alignment as binary, either…

Multimedia · Computer Science 2025-08-07 Ali Vosoughi , Jing Bi , Pinxin Liu , Yunlong Tang , Chenliang Xu

Video-quality measurement is a critical task in video processing. Nowadays, many implementations of new encoding standards - such as AV1, VVC, and LCEVC - use deep-learning-based decoding algorithms with perceptual metrics that serve as…

Computer Vision and Pattern Recognition · Computer Science 2023-02-08 Anastasia Antsiferova , Sergey Lavrushkin , Maksim Smirnov , Alexander Gushchin , Dmitriy Vatolin , Dmitriy Kulikov

Audio-based equipment condition monitoring suffers from a lack of standardized methodologies for algorithm selection, hindering reproducible research. This paper addresses this gap by introducing a comprehensive framework for the systematic…

Machine Learning · Computer Science 2026-03-20 Srijesh Pillai , Yodhin Agarwal , Zaheeruddin Ahmed

Our objective is to transform a video into a set of discrete audio-visual objects using self-supervised learning. To this end, we introduce a model that uses attention to localize and group sound sources, and optical flow to aggregate…

Computer Vision and Pattern Recognition · Computer Science 2020-08-11 Triantafyllos Afouras , Andrew Owens , Joon Son Chung , Andrew Zisserman

In industry, machine anomalous sound detection (ASD) is in great demand. However, collecting enough abnormal samples is difficult due to the high cost, which boosts the rapid development of unsupervised ASD algorithms. Autoencoder (AE)…

Sound · Computer Science 2023-11-16 Yifan Zhou , Dongxing Xu , Haoran Wei , Yanhua Long

Sense of hearing is crucial for autonomous vehicles (AVs) to better perceive its surrounding environment. Although visual sensors of an AV, such as camera, lidar, and radar, help to see its surrounding environment, an AV cannot see beyond…

Sound · Computer Science 2022-09-12 Finley Walden , Sagar Dasgupta , Mizanur Rahman , Mhafuzul Islam

Learning-based video quality assessment (VQA) has advanced rapidly, yet progress is increasingly constrained by a disconnect between model design and dataset curation. Model-centric approaches often iterate on fixed benchmarks, while…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Jian Zou , Xiaoyu Xu , Zhihua Wang , Yilin Wang , Balu Adsumilli , Kede Ma

In this paper, we present a novel deep fusion architecture for audio classification tasks. The multi-channel model presented is formed using deep convolution layers where different acoustic features are passed through each channel. To…

Sound · Computer Science 2018-11-05 Gaurav Bhatt , Akshita Gupta , Aditya Arora , Balasubramanian Raman

We propose a new framework for extracting visual information about a scene only using audio signals. Audio-based methods can overcome some of the limitations of vision-based methods i.e., they do not require "line-of-sight", are robust to…

Computer Vision and Pattern Recognition · Computer Science 2022-09-14 Fabrizio Pedersoli , Dryden Wiebe , Amin Banitalebi , Yong Zhang , George Tzanetakis , Kwang Moo Yi

We present a novel approach for interactive auditory object analysis with a humanoid robot. The robot elicits sensory information by physically shaking visually indistinguishable plastic capsules. It gathers the resulting audio signals from…

Robotics · Computer Science 2018-07-11 Manfred Eppe , Matthias Kerzel , Erik Strahl , Stefan Wermter

Noise reduction techniques based on deep learning have demonstrated impressive performance in enhancing the overall quality of recorded speech. While these approaches are highly performant, their application in audio engineering can be…

Sound · Computer Science 2023-10-18 Christian J. Steinmetz , Thomas Walther , Joshua D. Reiss

The deepfake generation of singing vocals is a concerning issue for artists in the music industry. In this work, we propose a singing voice deepfake detection (SVDD) system, which uses noise-variant encodings of open-AI's Whisper model. As…

Sound · Computer Science 2025-02-03 Falguni Sharma , Priyanka Gupta

State-of-the-art speaker verification models are based on deep learning techniques, which heavily depend on the handdesigned neural architectures from experts or engineers. We borrow the idea of neural architecture search(NAS) for the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-14 Xiaoyang Qu , Jianzong Wang , Jing Xiao

Recent advances in Audio-Visual Speech Recognition (AVSR) have led to unprecedented achievements in the field, improving the robustness of this type of system in adverse, noisy environments. In most cases, this task has been addressed…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Long (> 200 ms) audio inpainting, to recover a long missing part in an audio segment, could be widely applied to audio editing tasks and transmission loss recovery. It is a very challenging problem due to the high dimensional, complex and…

Sound · Computer Science 2019-11-18 Ya-Liang Chang , Kuan-Ying Lee , Po-Yu Wu , Hung-yi Lee , Winston Hsu

The performance of deep learning based edge detector has far exceeded that of humans, but the huge computational cost and complex training strategy hinder its further development and application. In this paper, we eliminate these…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Yachuan Li , Xavier Soria Pomab , Yongke Xi , Guanlin Li , Chaozhi Yang , Qian Xiao , Yun Bai , Zongmin LI

In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Haotian Wang , Yuxuan Xi , Hang Chen , Jun Du , Yan Song , Qing Wang , Hengshun Zhou , Chenxi Wang , Jiefeng Ma , Pengfei Hu , Ya Jiang , Shi Cheng , Jie Zhang , Yuzhe Weng