中文
相关论文

相关论文: PanopticQuery: Unified Query-Time Reasoning for 4D…

200 篇论文

Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. This…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Ying Zang , Xuanyi Liu , Yidong Han , Deyi Ji , Chaotao Ding , Yuanqi Hu , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu

Dominant paradigms for 4D LiDAR panoptic segmentation are usually required to train deep neural networks with large superimposed point clouds or design dedicated modules for instance association. However, these approaches perform redundant…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Gyeongrok Oh , Youngdong Jang , Jonghyun Choi , Suk-Ju Kang , Guang Lin , Sangpil Kim

Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks which couple powerful vision-language model (VLM) with…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yuxin Song , Wenkai Dong , Shizun Wang , Qi Zhang , Song Xue , Tao Yuan , Hu Yang , Haocheng Feng , Hang Zhou , Xinyan Xiao , Jingdong Wang

Open-vocabulary querying in 3D space is crucial for enabling more intelligent perception in applications such as robotics, autonomous systems, and augmented reality. However, most existing methods rely on 2D pixel-level parsing, leading to…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Yiren Lu , Yunlai Zhou , Yiran Qiao , Chaoda Song , Tuo Liang , Jing Ma , Huan Wang , Yu Yin

Previous surface reconstruction methods either suffer from low geometric accuracy or lengthy training times when dealing with real-world complex dynamic scenes involving multi-person activities, and human-object interactions. To tackle the…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Shuo Wang , Binbin Huang , Ruoyu Wang , Shenghua Gao

Modeling and understanding the 3D world is crucial for various applications, from augmented reality to robotic navigation. Recent advancements based on 3D Gaussian Splatting have integrated semantic information from multi-view images into…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Xingrui Wang , Cuiling Lan , Hanxin Zhu , Zhibo Chen , Yan Lu

High-dynamic scene reconstruction aims to represent static background with rigid spatial features and dynamic objects with deformed continuous spatiotemporal features. Typically, existing methods adopt unified representation model (e.g.,…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Hanyu Zhou , Haonan Wang , Haoyue Liu , Yuxing Duan , Luxin Yan , Gim Hee Lee

Understanding open-vocabulary 3D scenes with Gaussian-based representations remains challenging due to fragmented and spatially inconsistent semantic predictions across multi-view observations. In this paper, we present OpenGaFF, a novel…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Kunyi Li , Michael Niemeyer , Sen Wang , Stefano Gasperini , Nassir Navab , Federico Tombari

Despite significant advancements in dynamic neural rendering, existing methods fail to address the unique challenges posed by UAV-captured scenarios, particularly those involving monocular camera setups, top-down perspective, and multiple…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Jaehoon Choi , Dongki Jung , Christopher Maxey , Yonghan Lee , Sungmin Eum , Dinesh Manocha , Heesung Kwon

We consider the problem of novel-view synthesis (NVS) for dynamic scenes. Recent neural approaches have accomplished exceptional NVS results for static 3D scenes, but extensions to 4D time-varying scenes remain non-trivial. Prior efforts…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Yuanxing Duan , Fangyin Wei , Qiyu Dai , Yuhang He , Wenzheng Chen , Baoquan Chen

Most existing Dynamic Gaussian Splatting methods for complex dynamic urban scenarios rely on accurate object-level supervision from expensive manual labeling, limiting their scalability in real-world applications. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Su Sun , Cheng Zhao , Zhuoyang Sun , Yingjie Victor Chen , Mei Chen

Vision-centric occupancy networks, which represent the surrounding environment with uniform voxels with semantics, have become a new trend for safe driving of camera-only autonomous driving perception systems, as they are able to detect…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Yining Shi , Jiusi Li , Kun Jiang , Ke Wang , Yunlong Wang , Mengmeng Yang , Diange Yang

Understanding 3D scenes is pivotal for autonomous driving, robotics, and augmented reality. Recent semantic Gaussian Splatting approaches leverage large-scale 2D vision models to project 2D semantic features onto 3D scenes. However, they…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Tianyu Huang , Runnan Chen , Dongting Hu , Fengming Huang , Mingming Gong , Tongliang Liu

The reconstruction of dynamic 3D scenes using 3D Gaussian Splatting has shown significant promise. A key challenge, however, remains in modeling realistic motion, as most methods fail to align the motion of Gaussians with real-world…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Junoh Lee , Junmyeong Lee , Yeon-Ji Song , Inhwan Bae , Jisu Shin , Hae-Gon Jeon , Jin-Hwa Kim

We investigate whether video generative models can exhibit visuospatial intelligence, a capability central to human cognition, using only visual data. To this end, we present Video4Spatial, a framework showing that video diffusion models…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Zeqi Xiao , Yiwei Zhao , Lingxiao Li , Yushi Lan , Ning Yu , Rahul Garg , Roshni Cooper , Mohammad H. Taghavi , Xingang Pan

Open-world interactive object search in household environments requires understanding semantic relationships between objects and their surrounding context to guide exploration efficiently. Prior methods either rely on vision-language…

机器人学 · 计算机科学 2026-05-28 Imen Mahdi , Matteo Cassinelli , Fabien Despinoy , Tim Welschehold , Abhinav Valada

Perpetual 3D scene generation aims to produce long-range and coherent 3D view sequences, which is applicable for long-term video synthesis and 3D scene reconstruction. Existing methods follow a "navigate-and-imagine" fashion and rely on…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Chong Xia , Shengjun Zhang , Fangfu Liu , Chang Liu , Khodchaphun Hirunyaratsameewong , Yueqi Duan

3D occupancy prediction is critical for comprehensive scene understanding in vision-centric autonomous driving. Recent advances have explored utilizing 3D semantic Gaussians to model occupancy while reducing computational overhead, but they…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Xiaoyang Yan , Muleilan Pei , Shaojie Shen

Novel view synthesis of dynamic scenes has been an intriguing yet challenging problem. Despite recent advancements, simultaneously achieving high-resolution photorealistic results, real-time rendering, and compact storage remains a…

计算机视觉与模式识别 · 计算机科学 2024-04-08 Zhan Li , Zhang Chen , Zhong Li , Yi Xu

In real-world scenarios, environment changes caused by human or agent activities make it extremely challenging for robots to perform various long-term tasks. Recent works typically struggle to effectively understand and adapt to dynamic…

机器人学 · 计算机科学 2025-12-19 Luzhou Ge , Xiangyu Zhu , Zhuo Yang , Xuesong Li