English
Related papers

Related papers: FrameVGGT: Geometry-Aligned Frame-Level Memory for…

200 papers

Vision-Language Models (VLMs) have demonstrated strong performance on tasks such as video captioning and visual question answering. However, their growing scale and video-level inputs lead to significant computational and memory overhead,…

Hardware Architecture · Computer Science 2025-12-17 Chiyue Wei , Cong Guo , Junyao Zhang , Haoxuan Shan , Yifan Xu , Ziyue Zhang , Yudong Liu , Qinsi Wang , Changchun Zhou , Hai "Helen" Li , Yiran Chen

Real-world multimodal knowledge graphs (MKGs) are inherently heterogeneous, modeling entities that are associated with diverse modalities. Traditional knowledge graph embedding (KGE) methods excel at learning continuous representations of…

Artificial Intelligence · Computer Science 2026-03-16 Athanasios Efthymiou , Stevan Rudinac , Monika Kackovic , Nachoem Wijnberg , Marcel Worring

This work presents FG-Net, a general deep learning framework for large-scale point clouds understanding without voxelizations, which achieves accurate and real-time performance with a single NVIDIA GTX 1080 GPU. First, a novel noise and…

Computer Vision and Pattern Recognition · Computer Science 2021-06-18 Kangcheng Liu , Zhi Gao , Feng Lin , Ben M. Chen

Though recent advances in vision-language models (VLMs) have achieved remarkable progress across a wide range of multimodal tasks, understanding 3D spatial relationships from limited views remains a significant challenge. Previous reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Zhangquan Chen , Manyuan Zhang , Xinlei Yu , Xufang Luo , Mingze Sun , Zihao Pan , Xiang An , Yan Feng , Peng Pei , Xunliang Cai , Ruqi Huang

Current multi-view indoor 3D object detectors rely on sensor geometry that is costly to obtain (i.e., precisely calibrated multi-view camera poses) to fuse multi-view information into a global scene representation, limiting deployment in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yang Cao , Feize Wu , Dave Zhenyu Chen , Yingji Zhong , Lanqing Hong , Dan Xu

Lifelong embodied navigation requires agents to accumulate, retain, and exploit spatial-semantic experience across tasks, enabling efficient exploration in novel environments and rapid goal reaching in familiar ones. While object-centric…

Robotics · Computer Science 2025-12-29 Botao Ren , Junjun Hu , Xinda Xue , Minghua Luo , Jintao Chen , Haochen Bai , Liangliang You , Mu Xu

Autoregressive multimodal large language models (MLLMs) enable 3D generation but struggle to scale to high-resolution shapes due to inadequate 3D tokenizations. Compact set-based representations discard deterministic spatial ordering,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yuan Li , Congyi Zhang , Xifeng Gao , Xiaohu Guo

Recent text-to-video (T2V) generation methods have seen significant advancements. However, the majority of these works focus on producing short video clips of a single event (i.e., single-scene videos). Meanwhile, recent large language…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Han Lin , Abhay Zala , Jaemin Cho , Mohit Bansal

Vision agent memory has shown remarkable effectiveness in streaming video understanding. However, storing such memory for videos incurs substantial memory overhead, leading to high costs in both storage and computation. To address this…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Junxi Wang , Te Sun , Jiayi Zhu , Junxian Li , Haowen Xu , Zichen Wen , Xuming Hu , Zhiyu Li , Linfeng Zhang

Modern video generation frameworks based on Latent Diffusion Models suffer from inefficiencies in tokenization due to the Frame-Proportional Information Assumption. Existing tokenizers provide fixed temporal compression rates, causing the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Tianxiong Zhong , Xingye Tian , Boyuan Jiang , Xuebo Wang , Xin Tao , Pengfei Wan , Zhiwei Zhang

Autoregressive video synthesis offers a promising pathway for infinite-horizon generation but is fundamentally hindered by three intertwined challenges: semantic forgetting from context limitations, visual drift due to positional…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Jintao Chen , Chengyu Bai , Junjun Hu , Xinda Xue , Mu Xu

Instance-level object segmentation across disparate egocentric and exocentric views is a fundamental challenge in visual understanding, critical for applications in embodied AI and remote collaboration. This task is exceptionally difficult…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Yulu Gao , Bohao Zhang , Zongheng Tang , Jitong Liao , Wenjun Wu , Si Liu

Deep learning-based models have achieved remarkable performance in video super-resolution (VSR) in recent years, but most of these models are less applicable to online video applications. These methods solely consider the distortion quality…

Computer Vision and Pattern Recognition · Computer Science 2022-08-05 Jun Xiao , Xinyang Jiang , Ningxin Zheng , Huan Yang , Yifan Yang , Yuqing Yang , Dongsheng Li , Kin-Man Lam

Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. This…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Ying Zang , Xuanyi Liu , Yidong Han , Deyi Ji , Chaotao Ding , Yuanqi Hu , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu

Video retrieval is a challenging research topic bridging the vision and language areas and has attracted broad attention in recent years. Previous works have been devoted to representing videos by directly encoding from frame-level…

Computer Vision and Pattern Recognition · Computer Science 2020-06-17 Zerun Feng , Zhimin Zeng , Caili Guo , Zheng Li

Vision-language Navigation (VLN) requires an agent to understand visual observations and language instructions to navigate in unseen environments. Most existing approaches rely on static scene assumptions and struggle to generalize in…

Robotics · Computer Science 2026-03-24 Xiangchen Liu , Hanghan Zheng , Jeil Jeong , Minsung Yoon , Lin Zhao , Zhide Zhong , Haoang Li , Sung-Eui Yoon

Camera motion is a fundamental geometric signal that shapes visual perception and cinematic style, yet current video-capable vision-language models (VideoLLMs) rarely represent it explicitly and often fail on fine-grained motion primitives.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Haoan Feng , Sri Harsha Musunuri , Guan-Ming Su

This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-independent training-free approaches, VGGT-360 reformulates the task as panoramic reprojection…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Jiayi Yuan , Haobo Jiang , De Wen Soh , Na Zhao

The integration of visual inputs with large language models (LLMs) has led to remarkable advancements in multi-modal capabilities, giving rise to visual large language models (VLLMs). However, effectively harnessing VLLMs for intricate…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Renjie Pi , Lewei Yao , Jiahui Gao , Jipeng Zhang , Tong Zhang

High-quality 3D world models are pivotal for embodied intelligence and Artificial General Intelligence (AGI), underpinning applications such as AR/VR content creation and robotic navigation. Despite the established strong imaginative…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Yixiang Dai , Fan Jiang , Chiyu Wang , Mu Xu , Yonggang Qi
‹ Prev 1 8 9 10 Next ›