中文
相关论文

相关论文: CoMa: Contextual Massing Generation with Vision-La…

200 篇论文

Reasoning about images with rich text, such as charts and documents, is a critical application of vision-language models (VLMs). However, VLMs often struggle in these domains due to the scarcity of diverse text-rich vision-language data. To…

The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning. A recent line of work explores learning spatial reasoning directly from multi-view images,…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Kanghee Lee , Injae Lee , Minseok Kwak , Jungi Hong , Kwonyoung Ryu , Jaesik Park

Autonomous driving, particularly navigating complex and unanticipated scenarios, demands sophisticated reasoning and planning capabilities. While Multi-modal Large Language Models (MLLMs) offer a promising avenue for this, their use has…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Hidehisa Arai , Keita Miwa , Kento Sasaki , Yu Yamaguchi , Kohei Watanabe , Shunsuke Aoki , Issei Yamamoto

3D human motion generation has seen substantial advancement in recent years. While state-of-the-art approaches have improved performance significantly, they still struggle with complex and detailed motions unseen in training data, largely…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Shanlin Sun , Gabriel De Araujo , Jiaqi Xu , Shenghan Zhou , Hanwen Zhang , Ziheng Huang , Chenyu You , Xiaohui Xie

Vision-language models (VLMs) achieve remarkable success in single-image tasks. However, real-world scenarios often involve intricate multi-image inputs, leading to a notable performance decline as models struggle to disentangle critical…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Juntian Zhang , Chuanqi cheng , Yuhan Liu , Wei Liu , Jian Luan , Rui Yan

Interleaved image-text generation has emerged as a crucial multimodal task, aiming at creating sequences of interleaved visual and textual content given a query. Despite notable advancements in recent multimodal large language models…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Wei Chen , Lin Li , Yongqi Yang , Bin Wen , Fan Yang , Tingting Gao , Yu Wu , Long Chen

Vision Language Models (VLMs) perform well on standard video tasks but struggle with physics-related reasoning involving motion dynamics and spatial interactions. We present a novel approach to address this gap by translating physical-world…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Xiyang Wu , Zongxia Li , Jihui Jin , Guangyao Shi , Gouthaman KV , Vishnu Raj , Nilotpal Sinha , Jingxi Chen , Fan Du , Dinesh Manocha

Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usually project the image-based vision features to 3D space and learn the geometric information…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Yubo Cui , Zhiheng Li , Jiaqiang Wang , Zheng Fang

Creating digital models using Computer Aided Design (CAD) is a process that requires in-depth expertise. In industrial product development, this process typically involves entire teams of engineers, spanning requirements engineering, CAD…

人工智能 · 计算机科学 2025-03-07 Felix Ocker , Stefan Menzel , Ahmed Sadik , Thiago Rios

Vision-Language Models (VLMs) have recently witnessed significant progress in visual comprehension. As the permitting length of image context grows, VLMs can now comprehend a broader range of views and spaces. Current benchmarks provide…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Yiqi Zhu , Ziyue Wang , Can Zhang , Peng Li , Yang Liu

Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, VLM in the remote…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Xingxing Weng , Chao Pang , Gui-Song Xia

Generating enough and diverse data through augmentation offers an efficient solution to the time-consuming and labour-intensive process of collecting and annotating pixel-wise images. Traditional data augmentation techniques often face…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Jiaojiao Ye , Jiaxing Zhong , Qian Xie , Yuzhou Zhou , Niki Trigoni , Andrew Markham

Text conditioned generative models for images have yielded impressive results. Text conditioned floorplan generation as a special type of raster image generation task also received particular attention. However there are many use cases in…

计算与语言 · 计算机科学 2024-07-23 Zhi Hao Luo , Luis Lara , Ge Ya Luo , Florian Golemo , Christopher Beckham , Christopher Pal

Vision Language Models (VLMs) often struggle with chart understanding tasks, particularly in accurate chart description and complex reasoning. Synthetic data generation is a promising solution, while usually facing the challenge of noise…

人工智能 · 计算机科学 2025-08-19 Gongyao Jiang , Qiong Luo

Vision-Language Models (VLMs) have achieved remarkable success in various multi-modal tasks, but they are often bottlenecked by the limited context window and high computational cost of processing high-resolution image inputs and videos.…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Xubing Ye , Yukang Gan , Xiaoke Huang , Yixiao Ge , Yansong Tang

The development of web-based geospatial dashboards for risk analysis and decision support is often challenged by the difficulty in visualization of big, multi-dimensional environmental data, implementation complexity, and limited…

人机交互 · 计算机科学 2025-11-27 Haowen Xu , Jose Tupayachi , Xiao-Ying Yu

Effectively understanding urban scenes requires fine-grained spatial reasoning about objects, layouts, and depth cues. However, how well current vision-language models (VLMs), pretrained on general scenes, transfer these abilities to urban…

计算机视觉与模式识别 · 计算机科学 2025-09-01 Juneyoung Ro , Namwoo Kim , Yoonjin Yoon

We develop ImageNet-Think, a multimodal reasoning dataset designed to aid the development of Vision Language Models (VLMs) with explicit reasoning capabilities. Our dataset is built on 250,000 images from ImageNet21k dataset, providing…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Krishna Teja Chitty-Venkata , Murali Emani

Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image synthesis. However,…

计算与语言 · 计算机科学 2025-09-19 Pengyu Wang , Shaojun Zhou , Chenkun Tan , Xinghao Wang , Wei Huang , Zhen Ye , Zhaowei Li , Botian Jiang , Dong Zhang , Xipeng Qiu

Semantic location prediction from multimodal social media posts is a critical task with applications in personalized services and human mobility analysis. This paper introduces \textit{Contextualized Vision-Language Alignment (CoVLA)}, a…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Liu Jing , Amirul Rahman
‹ 上一页 1 2 3 10 下一页 ›