中文
相关论文

相关论文: Accurate and Efficient Stereo Matching via Attenti…

200 篇论文

The complementary characteristics of active and passive depth sensing techniques motivate the fusion of the Li-DAR sensor and stereo camera for improved depth perception. Instead of directly fusing estimated depths across LiDAR and stereo…

计算机视觉与模式识别 · 计算机科学 2019-04-08 Tsun-Hsuan Wang , Hou-Ning Hu , Chieh Hubert Lin , Yi-Hsuan Tsai , Wei-Chen Chiu , Min Sun

The Vision Transformer (ViT) has gained prominence for its superior relational modeling prowess. However, its global attention mechanism's quadratic complexity poses substantial computational burdens. A common remedy spatially groups tokens…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Qihang Fan , Huaibo Huang , Mingrui Chen , Ran He

Video saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip. Existing VSD-related works have mainly relied on the visual system but paid less attention to the audio aspect, while,…

计算机视觉与模式识别 · 计算机科学 2022-06-28 Chenglizhao Chen , Mengke Song , Wenfeng Song , Li Guo , Muwei Jian

State-of-the-art stereo matching methods typically use costly 3D convolutions to aggregate a full cost volume, but their computational demands make mobile deployment challenging. Directly applying 2D convolutions for cost aggregation often…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Gangwei Xu , Jiaxin Liu , Xianqi Wang , Junda Cheng , Yong Deng , Jinliang Zang , Yurui Chen , Xin Yang

The cost volume, capturing the similarity of possible correspondences across two input images, is a key ingredient in state-of-the-art optical flow approaches. When sampling correspondences to build the cost volume, a large neighborhood…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Huaizu Jiang , Erik Learned-Miller

In this paper, we propose ACA-Net, a lightweight, global context-aware speaker embedding extractor for Speaker Verification (SV) that improves upon existing work by using Asymmetric Cross Attention (ACA) to replace temporal pooling. ACA is…

In this study, we identify the inefficient attention phenomena in Large Vision-Language Models (LVLMs), notably within prominent models like LLaVA-1.5, QwenVL-Chat and Video-LLaVA. We find out that the attention computation over visual…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Liang Chen , Haozhe Zhao , Tianyu Liu , Shuai Bai , Junyang Lin , Chang Zhou , Baobao Chang

Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead and degrade inference efficiency. Prior studies typically…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Qiankun Ma , Ziyao Zhang , Haofei Wang , Jie Chen , Zhen Song , Hairong Zheng

In audio-visual navigation (AVN) tasks, an embodied agent must autonomously localize a sound source in unknown and complex 3D environments based on audio-visual signals. Existing methods often rely on static modality fusion strategies and…

人工智能 · 计算机科学 2025-09-23 Jia Li , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng

Audio-visual video parsing (AVVP) aims to detect event categories and their temporal boundaries in videos, typically under weak supervision. Existing methods mainly focus on (i) improving temporal modeling using attention-based…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Yaru Chen , Faegheh Sardari , Peiliang Zhang , Ruohao Guo , Yang Xiang , Zhenbo Li , Wenwu Wang

We introduce a novel cost aggregation network, dubbed Volumetric Aggregation with Transformers (VAT), to tackle the few-shot segmentation task by using both convolutions and transformers to efficiently handle high dimensional correlation…

计算机视觉与模式识别 · 计算机科学 2021-12-23 Sunghwan Hong , Seokju Cho , Jisu Nam , Seungryong Kim

The goal of Automatic Voice Over (AVO) is to generate speech in sync with a silent video given its text script. Recent AVO frameworks built upon text-to-speech synthesis (TTS) have shown impressive results. However, the current AVO learning…

音频与语音处理 · 电气工程与系统科学 2023-06-30 Junchen Lu , Berrak Sisman , Mingyang Zhang , Haizhou Li

Real-time stereo matching methods primarily focus on enhancing in-domain performance but often overlook the critical importance of generalization in real-world applications. In contrast, recent stereo foundation models leverage monocular…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Jiaxin Liu , Gangwei Xu , Xianqi Wang , Chengliang Zhang , Xin Yang

Deep neural networks have shown excellent performance for stereo matching. Many efforts focus on the feature extraction and similarity measurement of the matching cost computation step while less attention is paid on cost aggregation which…

计算机视觉与模式识别 · 计算机科学 2018-01-15 Lidong Yu , Yucheng Wang , Yuwei Wu , Yunde Jia

The recently proposed Conformer architecture has shown state-of-the-art performances in Automatic Speech Recognition by combining convolution with attention to model both local and global dependencies. In this paper, we study how to reduce…

音频与语音处理 · 电气工程与系统科学 2021-09-09 Maxime Burchi , Valentin Vielzeuf

Recent deep multi-view stereo (MVS) methods have widely incorporated transformers into cascade network for high-resolution depth estimation, achieving impressive results. However, existing transformer-based methods are constrained by their…

计算机视觉与模式识别 · 计算机科学 2024-02-05 Sicheng Wang , Hao Jiang , Lei Xiang

Contrastive learning has emerged as a powerful technique in audio-visual representation learning, leveraging the natural co-occurrence of audio and visual modalities in webscale video datasets. However, conventional contrastive audio-visual…

声音 · 计算机科学 2025-03-18 Ioannis Tsiamas , Santiago Pascual , Chunghsin Yeh , Joan Serrà

In this paper, we present a novel recurrent multi-view stereo network based on long short-term memory (LSTM) with adaptive aggregation, namely AA-RMVSNet. We firstly introduce an intra-view aggregation module to adaptively extract image…

计算机视觉与模式识别 · 计算机科学 2021-08-10 Zizhuang Wei , Qingtian Zhu , Chen Min , Yisong Chen , Guoping Wang

Audio-visual video segmentation~(AVVS) aims to generate pixel-level maps of sound-producing objects within image frames and ensure the maps faithfully adhere to the given audio, such as identifying and segmenting a singing person in a…

计算机视觉与模式识别 · 计算机科学 2023-09-21 Kexin Li , Zongxin Yang , Lei Chen , Yi Yang , Jun Xiao

Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy…

音频与语音处理 · 电气工程与系统科学 2019-05-02 George Sterpu , Christian Saam , Naomi Harte