解耦与去噪:应对视频时刻检索中的上下文错位
系统与控制
2024-08-15 v1 系统与控制
摘要
视频时刻检索旨在根据自然语言查询定位视频中的上下文时刻,是跨模态定位的一项基本任务。现有方法专注于增强所有时刻与文本描述之间的跨模态交互以实现视频理解。然而,由于时间线上语义分布不均和噪声视觉背景,持续与所有位置交互是不合理的。本文提出了一种跨模态上下文去噪网络(CDNet),通过解耦复杂关联和去除无关动态来实现精确的时刻检索。具体而言,我们提出了查询引导的语义解耦(QSD),通过估计全局和细粒度相关性来解耦视频时刻。我们提出了一种上下文感知的动态去噪(CDD),通过学习一组查询相关偏移来增强对对齐的时空细节的理解。在公共基准上的广泛实验表明,所提出的CDNet达到了最先进的性能。
引用
@article{arxiv.2408.07601,
title = {Microgrid Building Blocks for Dynamic Decoupling and Black Start Applications},
author = {Samrat Acharya and Priya Mana and Hisham Mahmood and Francis Tuffner and Alok Kumar Bharati},
journal= {arXiv preprint arXiv:2408.07601},
year = {2024}
}
备注
This paper is accepted for publication in IEEE PES Grid Edge Technologies Conference & Exposition 2025, San Diego, CA. The complete copyright version will be available on IEEE Xplore when the conference proceedings are published