English

TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos

Computer Vision and Pattern Recognition 2025-10-01 v1

Abstract

In this paper, we deal with the task of text-driven saliency detection in 360-degrees videos. For this, we introduce the TSV360 dataset which includes 16,000 triplets of ERP frames, textual descriptions of salient objects/events in these frames, and the associated ground-truth saliency maps. Following, we extend and adapt a SOTA visual-based approach for 360-degrees video saliency detection, and develop the TSalV360 method that takes into account a user-provided text description of the desired objects and/or events. This method leverages a SOTA vision-language model for data representation and integrates a similarity estimation module and a viewport spatio-temporal cross-attention mechanism, to discover dependencies between the different data modalities. Quantitative and qualitative evaluations using the TSV360 dataset, showed the competitiveness of TSalV360 compared to a SOTA visual-based approach and documented its competency to perform customized text-driven saliency detection in 360-degrees videos.

Keywords

Cite

@article{arxiv.2509.26208,
  title  = {TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos},
  author = {Ioannis Kontostathis and Evlampios Apostolidis and Vasileios Mezaris},
  journal= {arXiv preprint arXiv:2509.26208},
  year   = {2025}
}

Comments

IEEE CBMI 2025. This is the authors' accepted version. The final publication is available at https://ieeexplore.ieee.org/

R2 v1 2026-07-01T06:07:34.773Z