中文
相关论文

相关论文: CoSMo: A Multimodal Transformer for Page Stream Se…

200 篇论文

Self-supervised learning (SSL) has emerged as a promising paradigm for medical image analysis by harnessing unannotated data. Despite their potential, the existing SSL approaches overlook the high anatomical similarity inherent in medical…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Azad Singh , Deepak Mishra

Extreme Multimodal Summarization with Multimodal Output (XMSMO) becomes an attractive summarization approach by integrating various types of information to create extremely concise yet informative summaries for individual modalities.…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Sicheng Liu , Lintao Wang , Xiaogang Zhu , Xuequan Lu , Zhiyong Wang , Kun Hu

While Large Reasoning Models (LRMs) have demonstrated impressive capabilities in solving complex tasks through the generation of long reasoning chains, this reliance on verbose generation results in significant latency and computational…

计算与语言 · 计算机科学 2026-05-04 Runquan Gui , Jie Wang , Zhihai Wang , Chi Ma , Jianye Hao , Feng Wu

Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Yuriel Ryan , Rui Yang Tan , Kenny Tsu Wei Choo , Roy Ka-Wei Lee

We introduce Lumos, the first end-to-end multimodal question-answering system with text understanding capabilities. At the core of Lumos is a Scene Text Recognition (STR) component that extracts text from first person point-of-view images,…

Domain-generalized urban-scene semantic segmentation (USSS) aims to learn generalized semantic predictions across diverse urban-scene styles. Unlike domain gap challenges, USSS is unique in that the semantic categories are often similar in…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Qi Bi , Shaodi You , Theo Gevers

Advancements in text-to-image generative AI with large multimodal models are spreading into the field of image compression, creating high-quality representation of images at extremely low bit rates. This work introduces novel components to…

图像与视频处理 · 电气工程与系统科学 2025-06-02 Cheng-Lin Wu , Hyomin Choi , Ivan V. Bajić

We investigate the challenges of style transfer in multi-modal visual narratives. Among static visual narratives such as comics and manga, there are distinct visual styles in terms of presentation. They include style features across…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Yi-Chun Chen , Arnav Jhala

Recently, great success has been made in learning visual representations from text supervision, facilitating the emergence of text-supervised semantic segmentation. However, existing works focus on pixel grouping and cross-modal semantic…

计算机视觉与模式识别 · 计算机科学 2023-02-22 Pengzhen Ren , Changlin Li , Hang Xu , Yi Zhu , Guangrun Wang , Jianzhuang Liu , Xiaojun Chang , Xiaodan Liang

Semantic segmentation has made significant strides in pixel-level image understanding, yet it remains limited in capturing contextual and semantic relationships between objects. Current models, such as CNN and Transformer-based…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Ben Rahman

We propose the task of Panoptic Scene Completion (PSC) which extends the recently popular Semantic Scene Completion (SSC) task with instance-level information to produce a richer understanding of the 3D scene. Our PSC proposal utilizes a…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Anh-Quan Cao , Angela Dai , Raoul de Charette

Learning interpretable multimodal representations inherently relies on uncovering the conditional dependencies between heterogeneous features. However, sparse graph estimation techniques, such as Graphical Lasso (GLasso), to…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Fei Wang , Yutong Zhang , Xiong Wang

Inspired by the great success of language model (LM)-based pre-training, recent studies in visual document understanding have explored LM-based pre-training methods for modeling text within document images. Among them, pre-training that…

计算机视觉与模式识别 · 计算机科学 2023-09-25 Daehee Kim , Yoonsik Kim , DongHyun Kim , Yumin Lim , Geewook Kim , Taeho Kil

Understanding long-context visual information remains a fundamental challenge for vision-language models, particularly in agentic tasks such as GUI control and web navigation. While web pages and GUI environments are inherently structured…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Gyubeum Lim , Yemo Koo , Vijay Krishna Madisetti

With the novel and fast advances in the area of deep neural networks, several challenging image-based tasks have been recently approached by researchers in pattern recognition and computer vision. In this paper, we address one of these…

计算机视觉与模式识别 · 计算机科学 2022-11-11 Jônatas Wehrmann , Anderson Mattjie , Rodrigo C. Barros

Text-Based Person Search (TBPS) aims to retrieve target person images from a large-scale gallery using natural language descriptions, posing fundamental challenges in cross-modal representation learning. Existing methods often struggle to…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Jing Liu , Donglai Wei , Yang Liu , Sipeng Zhang , Tong Yang , Wei Zhou , Weiping Ding , Victor C. M. Leung

Page-level analysis of documents has been a topic of interest in digitization efforts, and multimodal approaches have been applied to both classification and page stream segmentation. In this work, we focus on capturing finer semantic…

机器学习 · 计算机科学 2022-05-27 Mehmet Arif Demirtaş , Berke Oral , Mehmet Yasin Akpınar , Onur Deniz

Multimodal manga analysis focuses on enhancing manga understanding with visual and textual features, which has attracted considerable attention from both natural language processing and computer vision communities. Currently, most comics…

计算与语言 · 计算机科学 2023-10-27 Hongcheng Guo , Boyang Wang , Jiaqi Bai , Jiaheng Liu , Jian Yang , Zhoujun Li

Real-time decoding of neural activity is central to neuroscience and neurotechnology applications, from closed-loop experiments to brain-computer interfaces, where models are subject to strict latency constraints. Traditional methods,…

神经元与认知 · 定量生物学 2025-11-10 Avery Hee-Woon Ryoo , Nanda H. Krishna , Ximeng Mao , Mehdi Azabou , Eva L. Dyer , Matthew G. Perich , Guillaume Lajoie

A system that enables blind or visually impaired users to access comics/manga would introduce a new medium of storytelling to this community. However, no such system currently exists. Generative vision-language models (VLMs) have shown…