中文
相关论文

相关论文: GeoFocus: Blending Efficient Global-to-Local Perce…

200 篇论文

Multimodal large language models (MLLMs) have achieved significant progress in image and language tasks due to the strong reasoning capability of large language models (LLMs). Nevertheless, most MLLMs suffer from limited spatial reasoning…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Jiajie Guo , Qingpeng Zhu , Jin Zeng , Xiaolong Wu , Changyong He , Weida Wang

Point cloud processing is a challenging task due to its sparsity and irregularity. Prior works introduce delicate designs on either local feature aggregator or global geometric architecture, but few combine both advantages. We propose…

计算机视觉与模式识别 · 计算机科学 2022-05-17 Renrui Zhang , Ziyao Zeng , Ziyu Guo , Xinben Gao , Kexue Fu , Jianbo Shi

Multimodal learning has rapidly advanced visual understanding, largely via multimodal large language models (MLLMs) that use powerful LLMs as cognitive cores. In visual generation, however, these powerful core models are typically reduced…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Han Lin , Xichen Pan , Ziqi Huang , Ji Hou , Jialiang Wang , Weifeng Chen , Zecheng He , Felix Juefei-Xu , Junzhe Sun , Zhipeng Fan , Ali Thabet , Mohit Bansal , Chu Wang

Geometric navigation is nowadays a well-established field of robotics and the research focus is shifting towards higher-level scene understanding, such as Semantic Mapping. When a robot needs to interact with its environment, it must be…

机器人学 · 计算机科学 2023-11-23 Federico Rollo , Gennaro Raiola , Andrea Zunino , Nikolaos Tsagarakis , Arash Ajoudani

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Jingzhi Li , Changjiang Luo , Ruoyu Chen , Hua Zhang , Wenqi Ren , Jianhou Gan , Xiaochun Cao

The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data interpretation. These models combine various data modalities,…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Ziyue Huang , Hongxi Yan , Qiqi Zhan , Shuai Yang , Mingming Zhang , Chenkai Zhang , YiMing Lei , Zeming Liu , Qingjie Liu , Yunhong Wang

Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization (CVGL), which aims to acquire the geolocation of images by image retrieval. To further…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Wei Wang , Dou Quan , Ning Huyan , Shuang Wang , Yi Li , Pei He , Licheng Jiao

Learning latent representations that capture both semantic and spatial information is central to efficient spatio-semantic reasoning. However, many existing approaches rely on implicit latent structures combined with dense feature maps or…

计算机视觉与模式识别 · 计算机科学 2026-05-13 SeongMin Jin , Doo Seok Jeong

Two-view correspondence learning is a key task in computer vision, which aims to establish reliable matching relationships for applications such as camera pose estimation and 3D reconstruction. However, existing methods have limitations in…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Shuyuan Lin , Mengtin Lo , Haosheng Chen , Yanjie Liang , Qiangqiang Wu

As the Large Language Model (LLM) becomes increasingly important in various domains. However, the following challenges still remain unsolved in accelerating LLM inference: (1) Synchronized partial softmax update. The softmax operation…

机器学习 · 计算机科学 2024-01-08 Ke Hong , Guohao Dai , Jiaming Xu , Qiuli Mao , Xiuhong Li , Jun Liu , Kangdi Chen , Yuhan Dong , Yu Wang

Remote sensing image scene classification remains a challenging task, primarily due to the complex spatial structures and multi-scale characteristics of ground objects. Although CNN-based methods excel at extracting local inductive biases,…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Yuanhao Tang , Xuechao Zou , Zhengpei Hu , Junliang Xing , Chengkun Zhang , Jianqiang Huang

Large Vision-Language Models (LVLMs) can accurately locate key objects in images, yet their attention to these objects tends to be very brief. Motivated by the hypothesis that sustained focus on key objects can improve LVLMs' visual…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Jianfei Zhao , Feng Zhang , Xin Sun , Chong Feng , Zhixing Tan

Geometric estimation is required for scene understanding and analysis in panoramic 360{\deg} images. Current methods usually predict a single feature, such as depth or surface normal. These methods can lack robustness, especially when…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Kun Huang , Fang-Lue Zhang , Fangfang Zhang , Yu-Kun Lai , Paul L. Rosin , Neil A. Dodgson

Multimodal large language models (MLLMs) represent images and video frames as visual tokens. Scaling from single images to hour-long videos, however, inflates the token budget far beyond practical limits. Popular pipelines therefore either…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Zirui Zhu , Hailun Xu , Yang Luo , Yong Liu , Kanchan Sarkar , Zhenheng Yang , Yang You

We present a novel local-global feature fusion framework for body-weight exercise recognition with floor-based dynamic pressure maps. One step further from the existing studies using deep neural networks mainly focusing on global feature…

计算机视觉与模式识别 · 计算机科学 2023-09-15 Davinder Pal Singh , Lala Shakti Swarup Ray , Bo Zhou , Sungho Suh , Paul Lukowicz

Global geolocation, which seeks to predict the geographical location of images captured anywhere in the world, is one of the most challenging tasks in the field of computer vision. In this paper, we introduce an innovative interactive…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Zhiyang Dou , Zipeng Wang , Xumeng Han , Guorong Li , Zhipei Huang , Zhenjun Han

In the geospatial domain, universal representation models are significantly less prevalent than their extensive use in natural language processing and computer vision. This discrepancy arises primarily from the high costs associated with…

人工智能 · 计算机科学 2024-12-19 Junlin He , Tong Nie , Wei Ma

Visual document understanding (VDU) has rapidly advanced with the development of powerful multi-modal language models. However, these models typically require extensive document pre-training data to learn intermediate representations and…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Souhail Bakkali , Sanket Biswas , Zuheng Ming , Mickaël Coustaty , Marçal Rusiñol , Oriol Ramos Terrades , Josep Lladós

The rapidly growing ecosystem of Large Language Models (LLMs) makes it increasingly challenging to manage and utilize the vast and dynamic pool of models effectively. We propose LOCUS, a method that produces low-dimensional vector…

机器学习 · 计算机科学 2026-01-30 Shivam Patel , William Cocke , Gauri Joshi

The emergence of large-scale pre-trained point cloud models has significantly advanced 3D scene understanding, but adapting these models to specific downstream tasks typically demands full fine-tuning, incurring high computational and…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Liyao Tang , Zhe Chen , Dacheng Tao