English

MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment

Computer Vision and Pattern Recognition 2026-07-01 v1 Machine Learning

Abstract

Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled representations. These challenges are severely exacerbated by two fundamental properties in the video domain: Temporal Misalignment, where textual descriptions often correlate only to specific, constrained temporal windows, leaving other frames text-irrelevant; and Semantic Asymmetry, which dictates a sparse, bidirectional, and non-equivalent relevance between frame-level visual details and caption-level concepts. This failure persists whether captions are short and temporally disjoint, creating ambiguity, or long and detailed, fostering entanglement between static objects and their temporal evolution. In this paper, we establish theoretical conditions that enable flexible alignment between video and text representations across the temporal dimension and at varying levels of granularity. Building on these theoretical insights, we introduce MoVA, Modular Long Video-Text Alignment, which learns dual asymmetric projections: a text-side projection that adaptively selects frame-aware subspaces of the caption, and a video-side projection that disentangles text-relevant visual concepts. Our framework ensures that the model can preserve global cross-modal semantics while disentangling evolving, frame-specific concepts and scale naturally to long captions and videos. Empirical evaluations show that MoVA outperforms existing methods in multiple video-text alignment tasks, demonstrating the effectiveness of our method.

Cite

@article{arxiv.2607.00858,
  title  = {MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment},
  author = {Peiyuan Zhu and Shaoan Xie and Zijian Li and Yifan Shen and Namrata Deka and Harsh Shrivastava and Guangyi Chen and Kun Zhang},
  journal= {arXiv preprint arXiv:2607.00858},
  year   = {2026}
}

Comments

ECCV 2026