This paper introduces CoSMo, a novel multimodal Transformer for Page Stream Segmentation (PSS) in comic books, a critical task for automated content understanding, as it is a necessary first stage for many downstream tasks like character analysis, story indexing, or metadata enrichment. We formalize PSS for this unique medium and curate a new 20,800-page annotated dataset. CoSMo, developed in vision-only and multimodal variants, consistently outperforms traditional baselines and significantly larger general-purpose vision-language models across F1-Macro, Panoptic Quality, and stream-level metrics. Our findings highlight the dominance of visual features for comic PSS macro-structure, yet demonstrate multimodal benefits in resolving challenging ambiguities. CoSMo establishes a new state-of-the-art, paving the way for scalable comic book analysis.
Cite
@article{arxiv.2507.10053,
title = {CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books},
author = {Marc Serra Ortega and Emanuele Vivoli and Artemis Llabrés and Dimosthenis Karatzas},
journal= {arXiv preprint arXiv:2507.10053},
year = {2025}
}