中文
相关论文

相关论文: Spatial-Aware Conditioned Fusion for Audio-Visual …

200 篇论文

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of…

计算机视觉与模式识别 · 计算机科学 2021-12-15 Yidi Li , Hong Liu , Hao Tang

Spoofing-robust speaker verification (SASV) combines the tasks of speaker and spoof detection to authenticate speakers under adversarial settings. Many SASV systems rely on fusion of speaker and spoof cues at embedding, score or decision…

音频与语音处理 · 电气工程与系统科学 2026-03-31 Oğuzhan Kurnaz , Jagabandhu Mishra , Tomi H. Kinnunen , Cemal Hanilçi

Remote sensing change detection plays a pivotal role in domains such as environmental monitoring, urban planning, and disaster assessment. However, existing methods typically rely on predefined categories and large-scale pixel-level…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Mingyu Dou , Shi Qiu , Ming Hu , Yifan Chen , Huping Ye , Xiaohan Liao , Zhe Sun

Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without retraining. Existing methods either require computationally…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Agata Żywot , Iason Skylitsis , Thijmen Nijdam , Zoe Tzifa-Kratira , Derck Prinzhorn , Konrad Szewczyk , Aritra Bhowmik

A novel approach to improving the performances of confocal scanning imaging is proposed. We experimentally demonstrate its feasibility using acoustic waves. It relies on a new way to encode spatial information using the temporal dimension.…

In this work, we focus on leveraging facial cues beyond the lip region for robust Audio-Visual Speech Enhancement (AVSE). The facial region, encompassing the lip region, reflects additional speech-related attributes such as gender, skin…

计算机视觉与模式识别 · 计算机科学 2023-11-27 Feixiang Wang , Shuang Yang , Shiguang Shan , Xilin Chen

In this study, we propose a novel multi-modal end-to-end neural approach for automated assessment of non-native English speakers' spontaneous speech using attention fusion. The pipeline employs Bi-directional Recurrent Convolutional Neural…

计算与语言 · 计算机科学 2021-11-30 Manraj Singh Grover , Yaman Kumar , Sumit Sarin , Payman Vafaee , Mika Hama , Rajiv Ratn Shah

Multi-label image classification is a critical task in machine learning that aims to accurately assign multiple labels to a single image. While existing methods often utilize attention mechanisms or graph convolutional networks to model…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Ren-Dong Xie , Zhi-Fen He , Bo Li , Bin Liu , Jin-Yan Hu

Hate speech in online videos is posing an increasingly serious threat to digital platforms, especially as video content becomes increasingly multimodal and context-dependent. Existing methods often struggle to effectively fuse the complex…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Shuonan Yang , Tailin Chen , Jiangbei Yue , Guangliang Cheng , Jianbo Jiao , Zeyu Fu

In this paper, we propose a solution for improving the quality of temporal sound localization. We employ a multimodal fusion approach to combine visual and audio features. High-quality visual features are extracted using a state-of-the-art…

声音 · 计算机科学 2024-07-03 Yurui Huang , Yang Yang , Shou Chen , Xiangyu Wu , Qingguo Chen , Jianfeng Lu

To maintain high perception performance among connected and autonomous vehicles (CAVs), in this paper, we propose an accuracy-aware and resource-efficient raw-level cooperative sensing and computing scheme among CAVs and road-side…

网络与互联网体系结构 · 计算机科学 2024-03-26 Xuehan Ye , Kaige Qu , Weihua Zhuang , Xuemin Shen

In this paper, we study the collaborative state fusion problem in a multi-agent environment, where mobile agents collaborate to track movable targets. Due to the limited sensing range and potential errors of on-board sensors, it is…

机器学习 · 计算机科学 2024-10-22 Tianlong Zhou , Jun Shang , Weixiong Rao

Correlation filters are special classifiers designed for shift-invariant object recognition, which are robust to pattern distortions. The recent literature shows that combining a set of sub-filters trained based on a single or a small group…

计算机视觉与模式识别 · 计算机科学 2018-02-14 Baochang Zhang , Shangzhen Luan , Chen Chen , Jungong Han , Wei Wang , Alessandro Perina , Ling Shao

This paper proposes a novel approach by integrating sensor fusion with deep reinforcement learning, specifically the Soft Actor-Critic (SAC) algorithm, to develop an optimal control policy for self-driving cars. Our system employs a…

系统与控制 · 电气工程与系统科学 2023-12-29 Amin Jalal Aghdasian , Amirhossein Heydarian Ardakani , Kianoush Aqabakee , Farzaneh Abdollahi

Temporal Action Localization (TAL) aims to identify actions' start, end, and class labels in untrimmed videos. While recent advancements using transformer networks and Feature Pyramid Networks (FPN) have enhanced visual feature recognition…

计算机视觉与模式识别 · 计算机科学 2023-10-06 Edward Fish , Jon Weinbren , Andrew Gilbert

This paper proposes a new task called spatial voice conversion, which aims to convert a target voice while preserving spatial information and non-target signals. Traditional voice conversion methods focus on single-channel waveforms,…

In rapidly-evolving domains such as autonomous driving, the use of multiple sensors with different modalities is crucial to ensure high operational precision and stability. To correctly exploit the provided information by each sensor in a…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Quentin Herau , Nathan Piasco , Moussab Bennehar , Luis Roldão , Dzmitry Tsishkou , Cyrille Migniot , Pascal Vasseur , Cédric Demonceaux

The images and sounds that we perceive undergo subtle but geometrically consistent changes as we rotate our heads. In this paper, we use these cues to solve a problem we call Sound Localization from Motion (SLfM): jointly estimating camera…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Ziyang Chen , Shengyi Qian , Andrew Owens

In smart cities, bandwidth-constrained Unmanned Aerial Vehicles (UAVs) often fail to relay mission-critical data in time, compromising real-time decision-making. This highlights the need for faster and more efficient transmission of only…

系统与控制 · 电气工程与系统科学 2026-01-15 Poorvi Joshi , Mohan Gurusamy

Keyword spotting (KWS) is crucial for many speech-driven applications, but robust KWS in noisy environments remains challenging. Conventional systems often rely on single-channel inputs and a cascaded pipeline separating front-end…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Rui Wang , Zhifei Zhang , Yu Gao , Xiaofeng Mou , Yi Xu