English
Related papers

Related papers: GroundFlow: A Plug-in Module for Temporal Reasonin…

200 papers

Visual grounding aims to localize the object referred to in an image based on a natural language query. Although progress has been made recently, accurately localizing target objects within multiple-instance distractions (multiple objects…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Minghang Zheng , Jiahua Zhang , Qingchao Chen , Yuxin Peng , Yang Liu

Vector quantization has emerged as a powerful tool in large-scale multimodal models, unifying heterogeneous representations through discrete token encoding. However, its effectiveness hinges on robust codebook design. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Hongxuan Li , Wencheng Zhu , Huiying Xu , Xinzhong Zhu , Pengfei Zhu

Long video understanding is challenging due to rich and complicated multimodal clues in long temporal range.Current methods adopt reasoning to improve the model's ability to analyze complex video clues in long videos via text-form…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Houlun Chen , Xin Wang , Guangyao Li , Yuwei Zhou , Yihan Chen , Jia Jia , Wenwu Zhu

Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VLM) approaches achieve promising results but rely on view-dependent reasoning or implicit…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Xuefei Sun , Xujia Zhang , Brendan Crowe , Doncey Albin , Christoffer Heckman

The recent advancement in video temporal grounding (VTG) has significantly enhanced fine-grained video understanding, primarily driven by multimodal large language models (MLLMs). With superior multimodal comprehension and reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Jianlong Wu , Wei Liu , Ye Liu , Meng Liu , Liqiang Nie , Zhouchen Lin , Chang Wen Chen

3D task planning has attracted increasing attention in human-robot interaction and embodied AI thanks to the recent advances in multimodal learning. However, most existing studies are facing two common challenges: 1) heavy reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Xueying Jiang , Wenhao Li , Xiaoqin Zhang , Ling Shao , Shijian Lu

Temporal Sentence Grounding in Videos (TSGV) aims to detect the event timestamps described by the natural language query from untrimmed videos. This paper discusses the challenge of achieving efficient computation in TSGV models while…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Renjie Liang , Yiming Yang , Hui Lu , Li Li

Inferring trajectories from longitudinal spatially-resolved omics data is fundamental to understanding the dynamics of structural and functional tissue changes in development, regeneration and repair, disease progression, and response to…

Machine Learning · Computer Science 2026-05-15 Santanu Subhash Rathod , Francesco Ceccarelli , Sean B. Holden , Pietro Liò , Xiao Zhang , Jovan Tanevski

Learning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual grounding approaches require a substantial number of bounding box…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Xiaoxu Xu , Yitian Yuan , Qiudan Zhang , Wenhui Wu , Zequn Jie , Lin Ma , Xu Wang

Deep neural network models have achieved remarkable progress in 3D scene understanding while trained in the closed-set setting and with full labels. However, the major bottleneck is that these models do not have the capacity to recognize…

Computer Vision and Pattern Recognition · Computer Science 2025-02-20 Kangcheng Liu , Yong-Jin Liu , Baoquan Chen

Despite the rapid progress of Multimodal Large Language Models (MLLMs), their ability to perform reliable visual grounding in high-stakes clinical software environments remains underexplored. Existing GUI benchmarks largely focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Rozain Shakeel , Abdul Rahman Mohammad Ali , Muneeb Mushtaq , Tausifa Jan Saleem , Tajamul Ashraf

The 3D visual grounding task aims to ground a natural language description to the targeted object in a 3D scene, which is usually represented in 3D point clouds. Previous works studied visual grounding under specific views. The…

Computer Vision and Pattern Recognition · Computer Science 2022-04-06 Shijia Huang , Yilun Chen , Jiaya Jia , Liwei Wang

This paper addresses the problem of 3D referring expression comprehension (REC) in autonomous driving scenario, which aims to ground a natural language to the targeted region in LiDAR point clouds. Previous approaches for REC usually focus…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Wenhao Cheng , Junbo Yin , Wei Li , Ruigang Yang , Jianbing Shen

Point cloud has drawn more and more research attention as well as real-world applications. However, many of these applications (e.g. autonomous driving and robotic manipulation) are actually based on sequential point clouds (i.e. four…

Computer Vision and Pattern Recognition · Computer Science 2022-04-22 Haiyan Wang , Yingli Tian

Establishing accurate point-to-point correspondences between non-rigid 3D shapes remains a critical challenge, particularly under non-isometric deformations and topological noise. Existing functional map pipelines suffer from ambiguities…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Tianwei Ye , Xiaoguang Mei , Yifan Xia , Fan Fan , Jun Huang , Jiayi Ma

Most existing Dynamic Gaussian Splatting methods for complex dynamic urban scenarios rely on accurate object-level supervision from expensive manual labeling, limiting their scalability in real-world applications. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Su Sun , Cheng Zhao , Zhuoyang Sun , Yingjie Victor Chen , Mei Chen

Currently, utilizing large language models to understand the 3D world is becoming popular. Yet existing 3D-aware LLMs act as black boxes: they output bounding boxes or textual answers without revealing how those decisions are made, and they…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Zhihao Yuan , Shuyi Jiang , Chun-Mei Feng , Yaolun Zhang , Shuguang Cui , Zhen Li , Na Zhao

3D Semantic Scene Completion (SSC) provides comprehensive scene geometry and semantics for autonomous driving perception, which is crucial for enabling accurate and reliable decision-making. However, existing SSC methods are limited to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Meng Wang , Fan Wu , Ruihui Li , Yunchuan Qin , Zhuo Tang , Kenli Li

In this technical report, we introduce our solution to human-centric spatio-temporal video grounding task. We propose a concise and effective framework named STVGFormer, which models spatiotemporal visual-linguistic dependencies with a…

Computer Vision and Pattern Recognition · Computer Science 2022-07-07 Zihang Lin , Chaolei Tan , Jian-Fang Hu , Zhi Jin , Tiancai Ye , Wei-Shi Zheng

Although 3D point cloud classification neural network models have been widely used, the in-depth interpretation of the activation of the neurons and layers is still a challenge. We propose a novel approach, named Relevance Flow, to…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Weiquan Liu , Minghao Liu , Shijun Zheng , Cheng Wang