中文
相关论文

相关论文: CAE-AV: Improving Audio-Visual Learning via Cross-…

200 篇论文

Joint understanding of video and language is an active research area with many applications. Prior work in this domain typically relies on learning text-video embeddings. One difficulty with this approach, however, is the lack of…

计算机视觉与模式识别 · 计算机科学 2020-01-17 Antoine Miech , Ivan Laptev , Josef Sivic

In this work, we focus on unsupervised vision-language-action mapping in the area of robotic manipulation. Recently, multiple approaches employing pre-trained large language and vision models have been proposed for this task. However, they…

机器人学 · 计算机科学 2025-05-29 Gabriela Sejnova , Michal Vavrecka , Karla Stepanova

Video caption refers to generating a descriptive sentence for a specific short video clip automatically, which has achieved remarkable success recently. However, most of the existing methods focus more on visual information while ignoring…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Wangli Hao , Zhaoxiang Zhang , He Guan , Guibo Zhu

Finetuning a pretrained vision model (PVM) is a common technique for learning downstream vision tasks. However, the conventional finetuning process with randomly sampled data points results in diminished training efficiency. To address this…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Rongyu Zhang , Zefan Cai , Huanrui Yang , Zidong Liu , Denis Gudovskiy , Tomoyuki Okuno , Yohei Nakata , Kurt Keutzer , Baobao Chang , Yuan Du , Li Du , Shanghang Zhang

Human speech processing is inherently multimodal, where visual cues (lip movements) help to better understand the speech in noise. Lip-reading driven speech enhancement significantly outperforms benchmark audio-only approaches at low…

计算机视觉与模式识别 · 计算机科学 2019-09-24 Ahsan Adeel , Mandar Gogate , Amir Hussain

Automatic emotion recognition (ER) has recently gained lot of interest due to its potential in many real-world applications. In this context, multimodal approaches have been shown to improve performance (over unimodal approaches) by…

计算机视觉与模式识别 · 计算机科学 2022-09-20 R Gnana Praveen , Eric Granger , Patrick Cardinal

Sound effect editing-modifying audio by adding, removing, or replacing elements-remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and…

多媒体 · 计算机科学 2025-11-27 Xinyue Guo , Xiaoran Yang , Lipan Zhang , Jianxuan Yang , Zhao Wang , Jian Luan

Automated audio captioning (AAC) has developed rapidly in recent years, involving acoustic signal processing and natural language processing to generate human-readable sentences for audio clips. The current models are generally based on the…

声音 · 计算机科学 2021-10-13 Zhongjie Ye , Helin Wang , Dongchao Yang , Yuexian Zou

The cochlear implant (CI) is a successful biomedical device that enables individuals with severe-to-profound hearing loss to perceive sound through electrical stimulation, yet listening in noise remains challenging. Recent deep learning…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Meng-Ping Lin , Enoch Hsin-Ho Huang , Shao-Yi Chien , Yu Tsao

Automated audio captioning (AAC), a task that mimics human perception as well as innovatively links audio processing and natural language processing, has overseen much progress over the last few years. AAC requires recognizing contents such…

声音 · 计算机科学 2023-11-17 Xuenan Xu , Zeyu Xie , Mengyue Wu , Kai Yu

Audio-visual learning has been a major pillar of multi-modal machine learning, where the community mostly focused on its modality-aligned setting, i.e., the audio and visual modality are both assumed to signal the prediction target. With…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Yung-Hsuan Lai , Yen-Chun Chen , Yu-Chiang Frank Wang

Robotic imitation learning has advanced from solving static tasks to addressing dynamic interaction scenarios, but testing and evaluation remain costly and challenging due to the need for real-time interaction with dynamic environments. We…

Embodied navigation demands comprehensive scene understanding and precise spatial reasoning. While image-text models excel at interpreting pixel-level color and lighting cues, 3D-text models capture volumetric structure and spatial…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Haihong Hao , Mingfei Han , Changlin Li , Zhihui Li , Xiaojun Chang

Audio-visual speech recognition (AVSR) has become critical for enhancing speech recognition in noisy environments by integrating both auditory and visual modalities. However, existing AVSR systems struggle to scale up without compromising…

音频与语音处理 · 电气工程与系统科学 2025-05-22 Sungnyun Kim , Kangwook Jang , Sangmin Bae , Sungwoo Cho , Se-Young Yun

In challenging real-life conditions such as extreme head-pose, occlusions, and low-resolution images where the visual information fails to estimate visual attention/gaze direction, audio signals could provide important and complementary…

计算机视觉与模式识别 · 计算机科学 2022-08-15 Shreya Ghosh , Abhinav Dhall , Munawar Hayat , Jarrod Knibbe

The integration of visual cues has revitalized the performance of the target speech extraction task, elevating it to the forefront of the field. Nevertheless, this multi-modal learning paradigm often encounters the challenge of modality…

声音 · 计算机科学 2024-05-07 Zhaoxi Mu , Xinyu Yang

Numerous studies have investigated the effectiveness of audio-visual multimodal learning for speech enhancement (AVSE) tasks, seeking a solution that uses visual data as auxiliary and complementary input to reduce the noise of noisy speech…

音频与语音处理 · 电气工程与系统科学 2022-02-02 Shang-Yi Chuang , Hsin-Min Wang , Yu Tsao

Audio-Visual Segmentation (AVS) aims to generate pixel-wise segmentation maps that correlate with the auditory signals of objects. This field has seen significant progress with numerous CNN and Transformer-based methods enhancing the…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Sitong Gong , Yunzhi Zhuge , Lu Zhang , Pingping Zhang , Huchuan Lu

Acoustic word embeddings (AWEs) aims to map a variable-length speech segment into a fixed-dimensional representation. High-quality AWEs should be invariant to variations, such as duration, pitch and speaker. In this paper, we introduce a…

音频与语音处理 · 电气工程与系统科学 2023-07-20 Jingru Lin , Xianghu Yue , Junyi Ao , Haizhou Li

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

多媒体 · 计算机科学 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller