中文
相关论文

相关论文: Interpretable Convolutional SyncNet

200 篇论文

Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model, and training that…

计算机视觉与模式识别 · 计算机科学 2024-01-30 Vladimir Iashin , Weidi Xie , Esa Rahtu , Andrew Zisserman

We present a new direction for increasing the interpretability of deep neural networks (DNNs) by promoting weight-input alignment during training. For this, we propose to replace the linear transforms in DNNs by our B-cos transform. As we…

计算机视觉与模式识别 · 计算机科学 2022-05-23 Moritz Böhle , Mario Fritz , Bernt Schiele

Infrared and visible image fusion targets to provide an informative image by combining complementary information from different sensors. Existing learning-based fusion approaches attempt to construct various loss functions to preserve…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Jinyuan Liu , Runjia Lin , Guanyao Wu , Risheng Liu , Zhongxuan Luo , Xin Fan

Recognizing symmetries in data allows for significant boosts in neural network training. In many cases, however, the underlying symmetry is present only in an idealized dataset, and is broken in the training data, due to effects such as…

高能物理 - 实验 · 物理学 2023-11-13 Edmund Witkowski , Daniel Whiteson

This paper investigates multimodal semantic non-orthogonal transmission and fusion in hybrid analog-digital massive multiple-input multiple-output (MIMO). A Transformer-based cross-modal source-channel semantic-aware network (CSC-SA-Net)…

信号处理 · 电气工程与系统科学 2025-12-15 Minghui Wu , Zhen Gao

Learning socially-aware motion representations is at the core of recent advances in multi-agent problems, such as human motion forecasting and robot navigation in crowds. Despite promising progress, existing representations learned with…

机器学习 · 计算机科学 2021-08-23 Yuejiang Liu , Qi Yan , Alexandre Alahi

Images captured by fisheye lenses violate the pinhole camera assumption and suffer from distortions. Rectification of fisheye images is therefore a crucial preprocessing step for many computer vision applications. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2018-04-16 Xiaoqing Yin , Xinchao Wang , Jun Yu , Maojun Zhang , Pascal Fua , Dacheng Tao

By extracting task-relevant information while maximally compressing the input, the information bottleneck (IB) principle has provided a guideline for learning effective and robust representations of the target inference. However, extending…

信息论 · 计算机科学 2025-01-22 Yuhan Wang , Youlong Wu , Shuai Ma , Ying-Jun Angela Zhang

In recent years, several unsupervised, "contrastive" learning algorithms in vision have been shown to learn representations that perform remarkably well on transfer tasks. We show that this family of algorithms maximizes a lower bound on…

机器学习 · 计算机科学 2020-06-08 Mike Wu , Chengxu Zhuang , Milan Mosse , Daniel Yamins , Noah Goodman

Semantic Change Detection (SCD) in remote sensing imagery requires accurately identifying land-cover changes across multi-temporal image pairs. Despite substantial advancements, including the introduction of transformer-based architectures,…

图像与视频处理 · 电气工程与系统科学 2025-11-11 Athulya Ratnayake , Buddhi Wijenayake , Praveen Sumanasekara , Roshan Godaliyadda , Vijitha Herath , Parakrama Ekanayake

Contour shape alignment is a fundamental but challenging problem in computer vision, especially when the observations are partial, noisy, and largely misaligned. Recent ConvNet-based architectures that were proposed to align image…

计算机视觉与模式识别 · 计算机科学 2020-05-26 VSR Veeravasarapu , Abhishek Goel , Deepak Mittal , Maneesh Singh

The perceptual loss has been widely used as an effective loss term in image synthesis tasks including image super-resolution, and style transfer. It was believed that the success lies in the high-level perceptual feature representations…

计算机视觉与模式识别 · 计算机科学 2021-03-22 Yifan Liu , Hao Chen , Yu Chen , Wei Yin , Chunhua Shen

In a typical sound event detection (SED) system, the existence of a sound event is detected at a frame level, and consecutive frames with the same event detected are combined as one sound event. The median filter is applied as a…

声音 · 计算机科学 2024-03-21 Tao Song

Today's deep learning systems deliver high performance based on end-to-end training. While they deliver strong performance, these systems are hard to interpret. To address this issue, we propose Semantic Bottleneck Networks (SBN): deep…

计算机视觉与模式识别 · 计算机科学 2019-07-30 Max Losch , Mario Fritz , Bernt Schiele

The challenges facing speech recognition systems, such as variations in pronunciations, adverse audio conditions, and the scarcity of labeled data, emphasize the necessity for a post-processing step that corrects recurring errors. Previous…

计算与语言 · 计算机科学 2023-10-18 Tomer Wullach , Shlomo E. Chazan

Transformation Synchronization is the problem of recovering absolute transformations from a given set of pairwise relative motions. Despite its usefulness, the problem remains challenging due to the influences from noisy and outlier…

计算机视觉与模式识别 · 计算机科学 2021-11-02 Zi Jian Yew , Gim Hee Lee

Binaural speech enhancement (BSE) aims to jointly improve the speech quality and intelligibility of noisy signals received by hearing devices and preserve the spatial cues of the target for natural listening. Existing methods often suffer…

声音 · 计算机科学 2025-01-09 Jingyuan Wang , Jie Zhang , Shihao Chen , Miao Sun

Contrastive learning has become a cornerstone of modern representation learning, allowing training with massive unlabeled data for both task-specific and general (foundation) models. A prototypical loss in contrastive training is InfoNCE…

机器学习 · 计算机科学 2026-03-02 Roy Betser , Eyal Gofer , Meir Yossef Levi , Guy Gilboa

Even though convolutional neural networks can classify objects in images very accurately, it is well known that the attention of the network may not always be on the semantically important regions of the scene. It has been observed that…

计算机视觉与模式识别 · 计算机科学 2022-02-10 Maliha Arif , Calvin Yong , Abhijit Mahalanobis

The photographs captured by digital cameras usually suffer from over or under exposure problems. For image exposure enhancement, the tasks of Single-Exposure Correction (SEC) and Multi-Exposure Fusion (MEF) are widely studied in the image…

图像与视频处理 · 电气工程与系统科学 2023-09-08 Jin Liang , Yuchen Yang , Anran Zhang , Jun Xu , Hui Li , Xiantong Zhen