English
Related papers

Related papers: CoSMo: A Multimodal Transformer for Page Stream Se…

200 papers

Self-supervised learning (SSL) has emerged as a promising paradigm for medical image analysis by harnessing unannotated data. Despite their potential, the existing SSL approaches overlook the high anatomical similarity inherent in medical…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Azad Singh , Deepak Mishra

Extreme Multimodal Summarization with Multimodal Output (XMSMO) becomes an attractive summarization approach by integrating various types of information to create extremely concise yet informative summaries for individual modalities.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Sicheng Liu , Lintao Wang , Xiaogang Zhu , Xuequan Lu , Zhiyong Wang , Kun Hu

While Large Reasoning Models (LRMs) have demonstrated impressive capabilities in solving complex tasks through the generation of long reasoning chains, this reliance on verbose generation results in significant latency and computational…

Computation and Language · Computer Science 2026-05-04 Runquan Gui , Jie Wang , Zhihai Wang , Chi Ma , Jianye Hao , Feng Wu

Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Yuriel Ryan , Rui Yang Tan , Kenny Tsu Wei Choo , Roy Ka-Wei Lee

We introduce Lumos, the first end-to-end multimodal question-answering system with text understanding capabilities. At the core of Lumos is a Scene Text Recognition (STR) component that extracts text from first person point-of-view images,…

Domain-generalized urban-scene semantic segmentation (USSS) aims to learn generalized semantic predictions across diverse urban-scene styles. Unlike domain gap challenges, USSS is unique in that the semantic categories are often similar in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Qi Bi , Shaodi You , Theo Gevers

Advancements in text-to-image generative AI with large multimodal models are spreading into the field of image compression, creating high-quality representation of images at extremely low bit rates. This work introduces novel components to…

Image and Video Processing · Electrical Eng. & Systems 2025-06-02 Cheng-Lin Wu , Hyomin Choi , Ivan V. Bajić

We investigate the challenges of style transfer in multi-modal visual narratives. Among static visual narratives such as comics and manga, there are distinct visual styles in terms of presentation. They include style features across…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Yi-Chun Chen , Arnav Jhala

Recently, great success has been made in learning visual representations from text supervision, facilitating the emergence of text-supervised semantic segmentation. However, existing works focus on pixel grouping and cross-modal semantic…

Computer Vision and Pattern Recognition · Computer Science 2023-02-22 Pengzhen Ren , Changlin Li , Hang Xu , Yi Zhu , Guangrun Wang , Jianzhuang Liu , Xiaojun Chang , Xiaodan Liang

Semantic segmentation has made significant strides in pixel-level image understanding, yet it remains limited in capturing contextual and semantic relationships between objects. Current models, such as CNN and Transformer-based…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Ben Rahman

We propose the task of Panoptic Scene Completion (PSC) which extends the recently popular Semantic Scene Completion (SSC) task with instance-level information to produce a richer understanding of the 3D scene. Our PSC proposal utilizes a…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Anh-Quan Cao , Angela Dai , Raoul de Charette

Learning interpretable multimodal representations inherently relies on uncovering the conditional dependencies between heterogeneous features. However, sparse graph estimation techniques, such as Graphical Lasso (GLasso), to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Fei Wang , Yutong Zhang , Xiong Wang

Inspired by the great success of language model (LM)-based pre-training, recent studies in visual document understanding have explored LM-based pre-training methods for modeling text within document images. Among them, pre-training that…

Computer Vision and Pattern Recognition · Computer Science 2023-09-25 Daehee Kim , Yoonsik Kim , DongHyun Kim , Yumin Lim , Geewook Kim , Taeho Kil

Understanding long-context visual information remains a fundamental challenge for vision-language models, particularly in agentic tasks such as GUI control and web navigation. While web pages and GUI environments are inherently structured…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Gyubeum Lim , Yemo Koo , Vijay Krishna Madisetti

With the novel and fast advances in the area of deep neural networks, several challenging image-based tasks have been recently approached by researchers in pattern recognition and computer vision. In this paper, we address one of these…

Computer Vision and Pattern Recognition · Computer Science 2022-11-11 Jônatas Wehrmann , Anderson Mattjie , Rodrigo C. Barros

Text-Based Person Search (TBPS) aims to retrieve target person images from a large-scale gallery using natural language descriptions, posing fundamental challenges in cross-modal representation learning. Existing methods often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Jing Liu , Donglai Wei , Yang Liu , Sipeng Zhang , Tong Yang , Wei Zhou , Weiping Ding , Victor C. M. Leung

Page-level analysis of documents has been a topic of interest in digitization efforts, and multimodal approaches have been applied to both classification and page stream segmentation. In this work, we focus on capturing finer semantic…

Machine Learning · Computer Science 2022-05-27 Mehmet Arif Demirtaş , Berke Oral , Mehmet Yasin Akpınar , Onur Deniz

Multimodal manga analysis focuses on enhancing manga understanding with visual and textual features, which has attracted considerable attention from both natural language processing and computer vision communities. Currently, most comics…

Computation and Language · Computer Science 2023-10-27 Hongcheng Guo , Boyang Wang , Jiaqi Bai , Jiaheng Liu , Jian Yang , Zhoujun Li

Real-time decoding of neural activity is central to neuroscience and neurotechnology applications, from closed-loop experiments to brain-computer interfaces, where models are subject to strict latency constraints. Traditional methods,…

Neurons and Cognition · Quantitative Biology 2025-11-10 Avery Hee-Woon Ryoo , Nanda H. Krishna , Ximeng Mao , Mehdi Azabou , Eva L. Dyer , Matthew G. Perich , Guillaume Lajoie

A system that enables blind or visually impaired users to access comics/manga would introduce a new medium of storytelling to this community. However, no such system currently exists. Generative vision-language models (VLMs) have shown…