中文
相关论文

相关论文: PixDLM: A Dual-Path Multimodal Language Model for …

200 篇论文

UAV vision-language navigation (VLN) requires an agent to navigate complex 3D environments from an egocentric perspective while following ambiguous multi-step instructions over long horizons. Existing zero-shot methods remain limited, as…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Dian Shao , Zhengzheng Xu , Peiyang Wang , Like Liu , Yule Wang , Jieqi Shi , Jing Huo

The escalating use of Unmanned Aerial Vehicles (UAVs) as remote sensing platforms has garnered considerable attention, proving invaluable for ground object recognition. While satellite remote sensing images face limitations in resolution…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Vlatko Spasev , Ivica Dimitrovski , Ivan Chorbev , Ivan Kitanovski

Surgical scene understanding is a cornerstone of computer-assisted intervention. While recent advances, particularly in surgical image segmentation, have driven progress, real-world clinical applications require a more holistic…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Jincai Huang , Shihao Zou , Yuchen Guo , Jingjing Li , Wei Ji , Kai Wang , Shanshan Wang , Weixin Si

As Vision-Language Models (VLMs) grow in sophistication, their ability to perform reasoning is coming under increasing supervision. While they excel at many tasks, their grasp of fundamental scientific principles, such as physics, remains…

Large vision-language models (VLMs) have garnered increasing interest in autonomous driving areas, due to their advanced capabilities in complex reasoning tasks essential for highly autonomous vehicle behavior. Despite their potential,…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Ming Nie , Renyuan Peng , Chunwei Wang , Xinyue Cai , Jianhua Han , Hang Xu , Li Zhang

With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still struggle with diverse applications such as robotics and autonomous driving. This paper aims to…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Peiwen Sun , Shiqiang Lang , Dongming Wu , Yi Ding , Kaituo Feng , Huadai Liu , Zhen Ye , Rui Liu , Yun-Hui Liu , Jianan Wang , Xiangyu Yue

Image Segmentation plays an essential role in computer vision and image processing with various applications from medical diagnosis to autonomous car driving. A lot of segmentation algorithms have been proposed for addressing specific…

计算机视觉与模式识别 · 计算机科学 2021-01-18 Yi Liu , Lutao Chu , Guowei Chen , Zewu Wu , Zeyu Chen , Baohua Lai , Yuying Hao

Vision-language models (VLMs) have emerged as a promising direction for end-to-end autonomous driving (AD) by jointly modeling visual observations, driving context, and language-based reasoning. However, existing VLM-based systems face a…

机器人学 · 计算机科学 2026-03-10 Ximeng Tao , Pardis Taghavi , Dimitar Filev , Reza Langari , Gaurav Pandey

Large multimodal models (LMMs) have achieved impressive progress in vision-language understanding, yet they face limitations in real-world applications requiring complex reasoning over a large number of images. Existing benchmarks for…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Jun Chen , Dannong Xu , Junjie Fei , Chun-Mei Feng , Mohamed Elhoseiny

The reasoning segmentation task, which demands a nuanced comprehension of intricate queries to accurately pinpoint object regions, is attracting increasing attention. However, Multi-modal Large Language Models (MLLM) often find it difficult…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Xiaoyi Bao , Siyang Sun , Shuailei Ma , Kecheng Zheng , Yuxin Guo , Guosheng Zhao , Yun Zheng , Xingang Wang

Comprehending 3D environments is vital for intelligent systems in domains like robotics and autonomous navigation. Voxel grids offer a structured representation of 3D space, but extracting high-level semantic meaning remains challenging.…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Alan Dao , Norapat Buppodom

This paper aims to achieve universal segmentation of arbitrary semantic level. Despite significant progress in recent years, specialist segmentation approaches are limited to specific tasks and data distribution. Retraining a new model for…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Yong Liu , Cairong Zhang , Yitong Wang , Jiahao Wang , Yujiu Yang , Yansong Tang

Synergistic spatial intelligence between UAVs and satellites is indispensable for emergency response and security operations, as it uniquely integrates macro-scale global coverage with dynamic, real-time local perception. However, the…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Dian Liu , Jie Feng , Di Li , Yuhui Zheng , Guanbin Li , Weisheng Dong , Guangming Shi

In the field of Vision-Language Navigation (VLN), aerial datasets remain limited in their ability to combine scale, diversity, and realism, often relying on either costly real-world scenes or visually limited simulations. To address these…

机器人学 · 计算机科学 2026-05-20 Jinhan Li , Xijie Huang , Zhaoqi Wang , Yijin Wang , Weiqi Ge , Qiyi He , Mo Zhu , Fei Gao , Yuze Wu , Xin Zhou

Though recent advances in vision-language models (VLMs) have achieved remarkable progress across a wide range of multimodal tasks, understanding 3D spatial relationships from limited views remains a significant challenge. Previous reasoning…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Zhangquan Chen , Manyuan Zhang , Xinlei Yu , Xufang Luo , Mingze Sun , Zihao Pan , Xiang An , Yan Feng , Peng Pei , Xunliang Cai , Ruqi Huang

The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Minkyu Kim , Sangheon Lee , Dongmin Park

Semantic segmentation from aerial views is a crucial task for autonomous drones, as they rely on precise and accurate segmentation to navigate safely and efficiently. However, aerial images present unique challenges such as diverse…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Benedikt Kolbeinsson , Krystian Mikolajczyk

Evaluating the performance of visual language models (VLMs) in graphic reasoning tasks has become an important research topic. However, VLMs still show obvious deficiencies in simulating human-level graphic reasoning capabilities,…

人工智能 · 计算机科学 2025-08-04 Jianyi Zhang , Xu Ji , Ziyin Zhou , Yuchen Zhou , Shubo Shi , Haoyu Wu , Zhen Li , Shizhao Liu

Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend these abilities by learning to perceive and reason, showing…

Remote sensing (RS) large vision-language models (LVLMs) have shown strong promise across visual grounding (VG) tasks. However, existing RS VG datasets predominantly rely on explicit referring expressions-such as relative position, relative…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Yue Zhou , Jue Chen , Zilun Zhang , Penghui Huang , Ran Ding , Zhentao Zou , PengFei Gao , Yuchen Wei , Ke Li , Xue Yang , Xue Jiang , Hongxin Yang , Jonathan Li