中文
相关论文

相关论文: Reliability-Aware Geometric Fusion for Robust Audi…

200 篇论文

We propose MAViD, a novel Multimodal framework for Audio-Visual Dialogue understanding and generation. Existing approaches primarily focus on non-interactive systems and are limited to producing constrained and unnatural human speech. The…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Youxin Pang , Jiajun Liu , Lingfeng Tan , Yong Zhang , Feng Gao , Xiang Deng , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

Underwater visual localization remains challenging due to wavelength-dependent attenuation, poor texture, and non-Gaussian sensor noise. We introduce MARVO, a physics-aware, learning-integrated odometry framework that fuses underwater image…

机器人学 · 计算机科学 2025-12-01 Sacchin Sundar , Atman Kikani , Aaliya Alam , Sumukh Shrote , A. Nayeemulla Khan , A. Shahina

Recent Audio-Visual Question Answering (AVQA) methods rely on complete visual and audio input to answer questions accurately. However, in real-world scenarios, issues such as device malfunctions and data transmission errors frequently…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Kyu Ri Park , Hong Joo Lee , Jung Uk Kim

Humans are adept at leveraging visual cues from lip movements for recognizing speech in adverse listening conditions. Audio-Visual Speech Recognition (AVSR) models follow similar approach to achieve robust speech recognition in noisy…

音频与语音处理 · 电气工程与系统科学 2024-05-24 Maxime Burchi , Krishna C. Puvvada , Jagadeesh Balam , Boris Ginsburg , Radu Timofte

Speech enhancement in audio-only settings remains challenging, particularly in the presence of interfering speakers. This paper presents a simple yet effective real-time audio-visual speech enhancement (AVSE) system, RAVEN, which isolates…

音频与语音处理 · 电气工程与系统科学 2025-08-05 T. Aleksandra Ma , Sile Yin , Li-Chia Yang , Shuo Zhang

This paper proposes a novel multimodal self-supervised architecture for energy-efficient audio-visual (AV) speech enhancement that integrates Graph Neural Networks with canonical correlation analysis (CCA-GNN). The proposed approach lays…

The accurate navigation of autonomous underwater vehicles critically depends on the precision of Doppler velocity log (DVL) velocity measurements. Recent advancements in deep learning have demonstrated significant potential in improving DVL…

机器人学 · 计算机科学 2025-12-16 Nadav Cohen , Itzik Klein

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. At each navigation step, the agent selects from possible candidate locations and then…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Zihan Wang , Xiangyang Li , Jiahao Yang , Yeqi Liu , Junjie Hu , Ming Jiang , Shuqiang Jiang

Advertisement (Ad) video violation detection is critical for ensuring platform compliance, but existing methods struggle with precise temporal grounding, noisy annotations, and limited generalization. We propose RAVEN, a novel framework…

计算与语言 · 计算机科学 2025-10-21 Deyi Ji , Yuekui Yang , Haiyang Wu , Shaoping Ma , Tianrun Chen , Lanyun Zhu

Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visual signals lack clear semantic correspondence. We propose to…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Seongah Kim , Dinh Phu Tran , Hyeontaek Hwang , Saad Wazir , Duc Do Minh , Daeyoung Kim

Unmanned Aerial Vehicle (UAV) Vision-and-Language Navigation (VLN) is vital for applications such as disaster response, logistics delivery, and urban inspection. However, existing methods often struggle with insufficient multimodal fusion,…

Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm to enhance large language models (LLMs) by conditioning generation on external evidence retrieved at inference time. While RAG addresses critical limitations of…

信息检索 · 计算机科学 2025-06-03 Chaitanya Sharma

Underwater Passive Acoustic Monitoring (UPAM) provides rich spatiotemporal data for long-term ecological analysis, but intrinsic noise and complex signal dependencies hinder model stability and generalization. Multilayered windowing has…

声音 · 计算机科学 2025-09-08 Nicholas R. Rasmussen , Rodrigue Rizk , Longwei Wang , KC Santosh

Retrieval-augmented generation (RAG)-based applications are gaining prominence due to their ability to leverage large language models (LLMs). These systems excel at combining retrieval mechanisms with generative capabilities, resulting in…

软件工程 · 计算机科学 2025-02-24 Rui Yang , Michael Fu , Chakkrit Tantithamthavorn , Chetan Arora , Lisa Vandenhurk , Joey Chua

Waveform-based deep learning faces a dilemma between nonparametric and parametric approaches. On one hand, convolutional neural networks (convnets) may approximate any linear time-invariant system; yet, in practice, their frequency…

声音 · 计算机科学 2024-07-09 Vincent Lostanlen , Daniel Haider , Han Han , Mathieu Lagrange , Peter Balazs , Martin Ehler

Delivering intelligent and adaptive navigation assistance in augmented reality (AR) requires more than visual cues, as it demands systems capable of interpreting flexible user intent and reasoning over both spatial and semantic context.…

人机交互 · 计算机科学 2025-08-26 Hsuan-Kung Yang , Tsu-Ching Hsiao , Ryoichiro Oka , Ryuya Nishino , Satoko Tofukuji , Norimasa Kobori

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Yuxin Guo , Shuailei Ma , Shijie Ma , Xiaoyi Bao , Chen-Wei Xie , Kecheng Zheng , Tingyu Weng , Siyang Sun , Yun Zheng , Wei Zou

Active Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails to capture…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yu Wang , Juhyung Ha , Frangil M. Ramirez , Yuchen Wang , David J. Crandall

Unmanned aerial vehicles (UAVs), commonly known as drones, are increasingly used across diverse domains, including logistics, agriculture, surveillance, and defense. While these systems provide numerous benefits, their misuse raises safety…

声音 · 计算机科学 2026-01-01 Rajdeep Chatterjee , Sudip Chakrabarty , Trishaani Acharjee , Deepanjali Mishra

Understanding the physical world requires perceptual models grounded in physical laws rather than mere statistical correlations. However, existing multimodal learning frameworks, focused on vision and language, lack physical consistency and…

人工智能 · 计算机科学 2025-11-26 Bo Pang , Chenxi Xu , Jierui Ren , Guoping Wang , Sheng Li