English
Related papers

Related papers: Interpretable Convolutional SyncNet

200 papers

Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model, and training that…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Vladimir Iashin , Weidi Xie , Esa Rahtu , Andrew Zisserman

We present a new direction for increasing the interpretability of deep neural networks (DNNs) by promoting weight-input alignment during training. For this, we propose to replace the linear transforms in DNNs by our B-cos transform. As we…

Computer Vision and Pattern Recognition · Computer Science 2022-05-23 Moritz Böhle , Mario Fritz , Bernt Schiele

Infrared and visible image fusion targets to provide an informative image by combining complementary information from different sensors. Existing learning-based fusion approaches attempt to construct various loss functions to preserve…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Jinyuan Liu , Runjia Lin , Guanyao Wu , Risheng Liu , Zhongxuan Luo , Xin Fan

Recognizing symmetries in data allows for significant boosts in neural network training. In many cases, however, the underlying symmetry is present only in an idealized dataset, and is broken in the training data, due to effects such as…

High Energy Physics - Experiment · Physics 2023-11-13 Edmund Witkowski , Daniel Whiteson

This paper investigates multimodal semantic non-orthogonal transmission and fusion in hybrid analog-digital massive multiple-input multiple-output (MIMO). A Transformer-based cross-modal source-channel semantic-aware network (CSC-SA-Net)…

Signal Processing · Electrical Eng. & Systems 2025-12-15 Minghui Wu , Zhen Gao

Learning socially-aware motion representations is at the core of recent advances in multi-agent problems, such as human motion forecasting and robot navigation in crowds. Despite promising progress, existing representations learned with…

Machine Learning · Computer Science 2021-08-23 Yuejiang Liu , Qi Yan , Alexandre Alahi

Images captured by fisheye lenses violate the pinhole camera assumption and suffer from distortions. Rectification of fisheye images is therefore a crucial preprocessing step for many computer vision applications. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2018-04-16 Xiaoqing Yin , Xinchao Wang , Jun Yu , Maojun Zhang , Pascal Fua , Dacheng Tao

By extracting task-relevant information while maximally compressing the input, the information bottleneck (IB) principle has provided a guideline for learning effective and robust representations of the target inference. However, extending…

Information Theory · Computer Science 2025-01-22 Yuhan Wang , Youlong Wu , Shuai Ma , Ying-Jun Angela Zhang

In recent years, several unsupervised, "contrastive" learning algorithms in vision have been shown to learn representations that perform remarkably well on transfer tasks. We show that this family of algorithms maximizes a lower bound on…

Machine Learning · Computer Science 2020-06-08 Mike Wu , Chengxu Zhuang , Milan Mosse , Daniel Yamins , Noah Goodman

Semantic Change Detection (SCD) in remote sensing imagery requires accurately identifying land-cover changes across multi-temporal image pairs. Despite substantial advancements, including the introduction of transformer-based architectures,…

Image and Video Processing · Electrical Eng. & Systems 2025-11-11 Athulya Ratnayake , Buddhi Wijenayake , Praveen Sumanasekara , Roshan Godaliyadda , Vijitha Herath , Parakrama Ekanayake

Contour shape alignment is a fundamental but challenging problem in computer vision, especially when the observations are partial, noisy, and largely misaligned. Recent ConvNet-based architectures that were proposed to align image…

Computer Vision and Pattern Recognition · Computer Science 2020-05-26 VSR Veeravasarapu , Abhishek Goel , Deepak Mittal , Maneesh Singh

The perceptual loss has been widely used as an effective loss term in image synthesis tasks including image super-resolution, and style transfer. It was believed that the success lies in the high-level perceptual feature representations…

Computer Vision and Pattern Recognition · Computer Science 2021-03-22 Yifan Liu , Hao Chen , Yu Chen , Wei Yin , Chunhua Shen

In a typical sound event detection (SED) system, the existence of a sound event is detected at a frame level, and consecutive frames with the same event detected are combined as one sound event. The median filter is applied as a…

Sound · Computer Science 2024-03-21 Tao Song

Today's deep learning systems deliver high performance based on end-to-end training. While they deliver strong performance, these systems are hard to interpret. To address this issue, we propose Semantic Bottleneck Networks (SBN): deep…

Computer Vision and Pattern Recognition · Computer Science 2019-07-30 Max Losch , Mario Fritz , Bernt Schiele

The challenges facing speech recognition systems, such as variations in pronunciations, adverse audio conditions, and the scarcity of labeled data, emphasize the necessity for a post-processing step that corrects recurring errors. Previous…

Computation and Language · Computer Science 2023-10-18 Tomer Wullach , Shlomo E. Chazan

Transformation Synchronization is the problem of recovering absolute transformations from a given set of pairwise relative motions. Despite its usefulness, the problem remains challenging due to the influences from noisy and outlier…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Zi Jian Yew , Gim Hee Lee

Binaural speech enhancement (BSE) aims to jointly improve the speech quality and intelligibility of noisy signals received by hearing devices and preserve the spatial cues of the target for natural listening. Existing methods often suffer…

Sound · Computer Science 2025-01-09 Jingyuan Wang , Jie Zhang , Shihao Chen , Miao Sun

Contrastive learning has become a cornerstone of modern representation learning, allowing training with massive unlabeled data for both task-specific and general (foundation) models. A prototypical loss in contrastive training is InfoNCE…

Machine Learning · Computer Science 2026-03-02 Roy Betser , Eyal Gofer , Meir Yossef Levi , Guy Gilboa

Even though convolutional neural networks can classify objects in images very accurately, it is well known that the attention of the network may not always be on the semantically important regions of the scene. It has been observed that…

Computer Vision and Pattern Recognition · Computer Science 2022-02-10 Maliha Arif , Calvin Yong , Abhijit Mahalanobis

The photographs captured by digital cameras usually suffer from over or under exposure problems. For image exposure enhancement, the tasks of Single-Exposure Correction (SEC) and Multi-Exposure Fusion (MEF) are widely studied in the image…

Image and Video Processing · Electrical Eng. & Systems 2023-09-08 Jin Liang , Yuchen Yang , Anran Zhang , Jun Xu , Hui Li , Xiantong Zhen