English
Related papers

Related papers: InterFormer: Interactive Local and Global Features…

200 papers

Sensor fusion approaches for intelligent self-driving agents remain key to driving scene understanding given visual global contexts acquired from input sensors. Specifically, for the local waypoint prediction task, single-modality networks…

Robotics · Computer Science 2024-02-01 Hwan-Soo Choi , Jongoh Jeong , Young Hoo Cho , Kuk-Jin Yoon , Jong-Hwan Kim

In this paper, we show that a simple self-supervised pre-trained audio model can achieve comparable inference efficiency to more complicated pre-trained models with speech transformer encoders. These speech transformers rely on mixing…

Sound · Computer Science 2024-02-09 Sungho Jeon , Ching-Feng Yeh , Hakan Inan , Wei-Ning Hsu , Rashi Rungta , Yashar Mehdad , Daniel Bikel

Conversational automatic speech recognition (ASR) is a task to recognize conversational speech including multiple speakers. Unlike sentence-level ASR, conversational ASR can naturally take advantages from specific characteristics of…

Sound · Computer Science 2022-02-18 Kun Wei , Yike Zhang , Sining Sun , Lei Xie , Long Ma

Driver action recognition, aiming to accurately identify drivers' behaviours, is crucial for enhancing driver-vehicle interactions and ensuring driving safety. Unlike general action recognition, drivers' environments are often challenging,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Ruoyu Wang , Wenqian Wang , Jianjun Gao , Dan Lin , Kim-Hui Yap , Bingbing Li

Micro-expressions are spontaneous, rapid and subtle facial movements that can neither be forged nor suppressed. They are very important nonverbal communication clues, but are transient and of low intensity thus difficult to recognize.…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Zhijun Zhai , Jianhui Zhao , Chengjiang Long , Wenju Xu , Shuangjiang He , Huijuan Zhao

Advanced visual localization techniques encompass image retrieval challenges and 6 Degree-of-Freedom (DoF) camera pose estimation, such as hierarchical localization. Thus, they must extract global and local features from input images.…

Computer Vision and Pattern Recognition · Computer Science 2022-12-27 Wenzheng Song , Ran Yan , Boshu Lei , Takayuki Okatani

Although numerous solutions have been proposed for image super-resolution, they are usually incompatible with low-power devices with many computational and memory constraints. In this paper, we address this problem by proposing a simple yet…

Computer Vision and Pattern Recognition · Computer Science 2023-02-28 Long Sun , Jiangxin Dong , Jinhui Tang , Jinshan Pan

Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-02 George Sterpu , Christian Saam , Naomi Harte

Face recognition in the infrared (IR) band has become an important supplement to visible light face recognition due to its advantages of independent background light, strong penetration, ability of imaging under harsh environments such as…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Zhicheng Cao , Jiaxuan Zhang , Liaojun Pang

A fine-grained understanding of egocentric human-environment interactions is crucial for developing next-generation embodied agents. One fundamental challenge in this area involves accurately parsing hands and active objects. While…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Yuejiao Su , Yi Wang , Lei Yao , Yawen Cui , Lap-Pui Chau

To let the state-of-the-art end-to-end ASR model enjoy data efficiency, as well as much more unpaired text data by multi-modal training, one needs to address two problems: 1) the synchronicity of feature sampling rates between speech and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-02 Yuhang Yang , Haihua Xu , Hao Huang , Eng Siong Chng , Sheng Li

Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving the recognition accuracy. However, the lack of high-quality…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-01 Jinming Chen , Jingyi Fang , Yuanzhong Zheng , Yaoxuan Wang , Haojun Fei

Automatic Speech Recognition (ASR) has advanced with Speech Foundation Models (SFMs), yet performance degrades on dysarthric speech due to variability and limited data. This study as part of the submission to the Speech Accessibility…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Alexandre Ducorroy , Rachid Riad

Neural networks for visual content understanding have recently evolved from convolutional ones (CNNs) to transformers. The prior (CNN) relies on small-windowed kernels to capture the regional clues, demonstrating solid local expressiveness.…

Computer Vision and Pattern Recognition · Computer Science 2023-08-17 Zixuan Su , Hao Zhang , Jingjing Chen , Lei Pang , Chong-Wah Ngo , Yu-Gang Jiang

While the pursuit of higher accuracy in deepfake detection remains a central goal, there is an increasing demand for precise localization of manipulated regions. Despite the remarkable progress made in classification-based detection,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Chao Shuai , Gaojian Wang , Kun Pan , Tong Wu , Fanli Jin , Haohan Tan , Mengxiang Li , Zhenguang Liu , Feng Lin , Kui Ren

Multimodal image fusion aims to integrate information from different imaging techniques to produce a comprehensive, detail-rich single image for downstream vision tasks. Existing methods based on local convolutional neural networks (CNNs)…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Xinyu Xie , Yawen Cui , Tao Tan , Xubin Zheng , Zitong Yu

End-to-end automatic speech recognition suffers from adaptation to unknown target domain speech despite being trained with a large amount of paired audio--text data. Recent studies estimate a linguistic bias of the model as the internal…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-16 Emiru Tsunoo , Yosuke Kashiwagi , Chaitanya Narisetty , Shinji Watanabe

Unsupervised video person re-identification (reID) methods usually depend on global-level features. And many supervised reID methods employed local-level features and achieved significant performance improvements. However, applying…

Computer Vision and Pattern Recognition · Computer Science 2022-02-15 Xianghao Zang , Ge Li , Wei Gao , Xiujun Shu

SqueezeFormer has recently shown impressive performance in automatic speech recognition (ASR). However, its inference speed suffers from the quadratic complexity of softmax-attention (SA). In addition, limited by the large convolution…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-16 Yuguang Yang , Yu Pan , Jingjing Yin , Jiangyu Han , Lei Ma , Heng Lu

Local Feature Matching, an essential component of several computer vision tasks (e.g., structure from motion and visual localization), has been effectively settled by Transformer-based methods. However, these methods only integrate…

Computer Vision and Pattern Recognition · Computer Science 2023-10-23 Xinyu Zhang , Li Wang , Zhiqiang Jiang , Kun Dai , Tao Xie , Lei Yang , Wenhao Yu , Yang Shen , Jun Li