English
Related papers

Related papers: PixelWorld: How Far Are We from Perceiving Everyth…

200 papers

Language models~(LMs) gradually become general-purpose interfaces in the interactive and embodied world, where the understanding of physical concepts is an essential prerequisite. However, it is not yet clear whether LMs can understand…

Computation and Language · Computer Science 2023-05-24 Lei Li , Jingjing Xu , Qingxiu Dong , Ce Zheng , Qi Liu , Lingpeng Kong , Xu Sun

We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a…

Computer Vision and Pattern Recognition · Computer Science 2023-08-04 Weiyun Wang , Min Shi , Qingyun Li , Wenhai Wang , Zhenhang Huang , Linjie Xing , Zhe Chen , Hao Li , Xizhou Zhu , Zhiguo Cao , Yushi Chen , Tong Lu , Jifeng Dai , Yu Qiao

On-screen game footage contains rich contextual information that players process when playing and experiencing a game. Learning pixel representations of games can benefit artificial intelligence across several downstream tasks including…

Computer Vision and Pattern Recognition · Computer Science 2023-07-24 Chintan Trivedi , Konstantinos Makantasis , Antonios Liapis , Georgios N. Yannakakis

Since the emergence of Vision Transformer (ViT), it has been widely used in generative language model and generative visual model. Especially in the current state-of-art open source multimodal models, ViT obtained by CLIP or SigLIP method…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Haun Leung , ZiNan Wang

In this paper, we address the semantic segmentation problem with a focus on the context aggregation strategy. Our motivation is that the label of a pixel is the category of the object that the pixel belongs to. We present a simple yet…

Computer Vision and Pattern Recognition · Computer Science 2021-05-03 Yuhui Yuan , Xiaokang Chen , Xilin Chen , Jingdong Wang

Document understanding and GUI interaction are among the highest-value applications of Vision-Language Models (VLMs), yet they impose exceptionally heavy computational burden: fine-grained text and small UI elements demand high-resolution…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Nan Wang , Zhiwei Jin , Chen Chen , Haonan Lu

The key to integrating visual language tasks is to establish a good alignment strategy. Recently, visual semantic representation has achieved fine-grained visual understanding by dividing grids or image patches. However, the coarse-grained…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Siyu Zhang , Yeming Chen , Yaoru Sun , Fang Wang , Jun Yang , Lizhi Bai , Shangce Gao

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Ali Athar , Xueqing Deng , Liang-Chieh Chen

Representing the semantics of words is a long-standing problem for the natural language processing community. Most methods compute word semantics given their textual context in large corpora. More recently, researchers attempted to…

Computation and Language · Computer Science 2017-11-10 Éloi Zablocki , Benjamin Piwowarski , Laure Soulier , Patrick Gallinari

Vision-Language Models (VLMs) have demonstrated remarkable performance across a variety of real-world tasks. However, existing VLMs typically process visual information by serializing images, a method that diverges significantly from the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Yueyan Li , Chenggong Zhao , Zeyuan Zang , Caixia Yuan , Xiaojie Wang

We introduce Perception Encoder (PE), a state-of-the-art vision encoder for image and video understanding trained via simple vision-language learning. Traditionally, vision encoders have relied on a variety of pretraining objectives, each…

Artificial intelligence seems to be taking over the world with systems that model pixels, words, and phonemes. The world is arguably made up, not of pixels, words, and phonemes but of entities (objects, things, including events) with…

Artificial Intelligence · Computer Science 2025-11-06 David Poole

Image descriptions can help visually impaired people to quickly understand the image content. While we made significant progress in automatically describing images and optical character recognition, current approaches are unable to include…

Computer Vision and Pattern Recognition · Computer Science 2020-08-05 Oleksii Sidorov , Ronghang Hu , Marcus Rohrbach , Amanpreet Singh

Scene understanding remains a significant challenge in the computer vision community. The visual psychophysics literature has demonstrated the importance of interdependence among parts of the scene. Yet, the majority of methods in computer…

Computer Vision and Pattern Recognition · Computer Science 2011-08-23 Jason J. Corso

While many image colorization algorithms have recently shown the capability of producing plausible color versions from gray-scale photographs, they still suffer from the problems of context confusion and edge color bleeding. To address…

Computer Vision and Pattern Recognition · Computer Science 2018-08-07 Jiaojiao Zhao , Li Liu , Cees G. M. Snoek , Jungong Han , Ling Shao

Text to image transformation for input to neural networks requires intermediate steps. This paper attempts to present a new approach to pixel normalization so as to convert textual data into image, suitable as input for neural networks.…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Parth Sane , Ravindra Agrawal

We introduce and demonstrate how to effectively train multilingual machine translation models with pixel representations. We experiment with two different data settings with a variety of language and script coverage, demonstrating improved…

Computation and Language · Computer Science 2023-10-25 Elizabeth Salesky , Neha Verma , Philipp Koehn , Matt Post

The rapid advancement of generative models has intensified the challenge of detecting and interpreting visual forgeries, necessitating robust frameworks for image forgery detection while providing reasoning as well as localization. While…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Ipsita Praharaj , Yukta Butala , Badrikanath Praharaj , Yash Butala

Chain-of-thought reasoning has significantly improved the performance of Large Language Models (LLMs) across various domains. However, this reasoning process has been confined exclusively to textual space, limiting its effectiveness in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Haozhe Wang , Alex Su , Weiming Ren , Fangzhen Lin , Wenhu Chen

With the rapid progression of deep learning technologies, multi-modality image fusion has become increasingly prevalent in object detection tasks. Despite its popularity, the inherent disparities in how different sources depict scene…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Xingyuan Li , Yang Zou , Jinyuan Liu , Zhiying Jiang , Long Ma , Xin Fan , Risheng Liu