中文
相关论文

相关论文: GSWorld: Closed-Loop Photo-Realistic Simulation Su…

200 篇论文

We introduce G-Style, a novel algorithm designed to transfer the style of an image onto a 3D scene represented using Gaussian Splatting. Gaussian Splatting is a powerful 3D representation for novel view synthesis, as -- compared to other…

图形学 · 计算机科学 2024-09-06 Áron Samuel Kovács , Pedro Hermosilla , Renata G. Raidou

The emergence of 3D Gaussian Splatting for fast and high-quality novel view synthesize has opened up the possibility to construct photo-realistic simulations from video for robotic reinforcement learning. While the approach has been…

机器人学 · 计算机科学 2024-10-28 Liyou Zhou , Oleg Sinavski , Athanasios Polydoros

The ability for robots to perform efficient and zero-shot grasping of object parts is crucial for practical applications and is becoming prevalent with recent advances in Vision-Language Models (VLMs). To bridge the 2D-to-3D gap for…

机器人学 · 计算机科学 2024-09-04 Mazeyu Ji , Ri-Zhao Qiu , Xueyan Zou , Xiaolong Wang

3D reconstruction of indoor and urban environments is a prominent research topic with various downstream applications. However, existing geometric priors for addressing low-texture regions in indoor and urban settings often lack global…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Xiyu Zhang , Chong Bao , Yipeng Chen , Hongjia Zhai , Yitong Dong , Hujun Bao , Zhaopeng Cui , Guofeng Zhang

Scene reconstruction has emerged as a central challenge in computer vision, with approaches such as Neural Radiance Fields (NeRF) and Gaussian Splatting achieving remarkable progress. While Gaussian Splatting demonstrates strong performance…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Alexander Valverde , Brian Xu , Yuyin Zhou , Meng Xu , Hongyun Wang

Dynamic scene reconstruction has garnered significant attention in recent years due to its capabilities in high-quality and real-time rendering. Among various methodologies, constructing a 4D spatial-temporal representation, such as 4D-GS,…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Weiwei Cai , Weicai Ye , Peng Ye , Tong He , Tao Chen

Visual localization is the task of estimating a camera pose in a known environment. In this paper, we utilize 3D Gaussian Splatting (3DGS)-based representations for accurate and privacy-preserving visual localization. We propose Gaussian…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Maxime Pietrantoni , Gabriela Csurka , Torsten Sattler

We present Real-time Gaussian SLAM (RTG-SLAM), a real-time 3D reconstruction system with an RGBD camera for large-scale environments using Gaussian splatting. The system features a compact Gaussian representation and a highly efficient…

计算机视觉与模式识别 · 计算机科学 2024-05-10 Zhexi Peng , Tianjia Shao , Yong Liu , Jingke Zhou , Yin Yang , Jingdong Wang , Kun Zhou

3D Gaussian Splatting (3DGS) has recently gained popularity for efficient scene rendering by representing scenes as explicit sets of anisotropic 3D Gaussians. However, most existing work focuses primarily on modeling external surfaces. In…

图像与视频处理 · 电气工程与系统科学 2026-01-12 Shuxin Liang , Yihan Xiao , Wenlu Tang

We propose ActiveSplat, an autonomous high-fidelity reconstruction system leveraging Gaussian splatting. Taking advantage of efficient and realistic rendering, the system establishes a unified framework for online mapping, viewpoint…

机器人学 · 计算机科学 2025-06-17 Yuetao Li , Zijia Kuang , Ting Li , Qun Hao , Zike Yan , Guyue Zhou , Shaohui Zhang

We propose a dense RGBD SLAM system based on 3D Gaussian Splatting that provides metrically accurate pose tracking and visually realistic reconstruction. To this end, we first propose a Gaussian densification strategy based on the rendering…

机器人学 · 计算机科学 2024-10-03 Shuo Sun , Malcolm Mielle , Achim J. Lilienthal , Martin Magnusson

Simultaneous Localization and Mapping (SLAM) is one of the most important environment-perception and navigation algorithms for computer vision, robotics, and autonomous cars/drones. Hence, high quality and fast mapping becomes a fundamental…

Tracking the 6DoF pose of unknown objects in monocular RGB video sequences is crucial for robotic manipulation. However, existing approaches typically rely on accurate depth information, which is non-trivial to obtain in real-world…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Zhiyuan Chen , Fan Lu , Guo Yu , Bin Li , Sanqing Qu , Yuan Huang , Changhong Fu , Guang Chen

We introduce GE-Sim 2.0 (Genie Envisioner World Simulator 2.0), a closed-loop video world simulator for robotic manipulation. Building on the action-conditioned video generation framework of Genie Envisioner, GE-Sim 2.0 is re-trained on…

Text-based generation and editing of 3D scenes hold significant potential for streamlining content creation through intuitive user interactions. While recent advances leverage 3D Gaussian Splatting (3DGS) for high-fidelity and real-time…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Hyojun Go , Byeongjun Park , Jiho Jang , Jin-Young Kim , Soonwoo Kwon , Changick Kim

High-quality scene reconstruction and novel view synthesis based on Gaussian Splatting (3DGS) typically require steady, high-quality photographs, often impractical to capture with handheld cameras. We present a method that adapts to camera…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Otto Seiskari , Jerry Ylilammi , Valtteri Kaatrasalo , Pekka Rantalankila , Matias Turkulainen , Juho Kannala , Esa Rahtu , Arno Solin

Multi-task robotic bimanual manipulation is becoming increasingly popular as it enables sophisticated tasks that require diverse dual-arm collaboration patterns. Compared to unimanual manipulation, bimanual tasks pose challenges to…

机器人学 · 计算机科学 2025-06-25 Tengbo Yu , Guanxing Lu , Zaijia Yang , Haoyuan Deng , Season Si Chen , Jiwen Lu , Wenbo Ding , Guoqiang Hu , Yansong Tang , Ziwei Wang

We propose a novel 3D deepfake generation framework based on 3D Gaussian Splatting that enables realistic, identity-preserving face swapping and reenactment in a fully controllable 3D space. Compared to conventional 2D deepfake approaches…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Wending Liu , Siyun Liang , Huy H. Nguyen , Isao Echizen

Accurate scene perception is critical for vision-based robotic manipulation. Existing approaches typically follow either a Vision-to-Action (V-A) paradigm, predicting actions directly from visual inputs, or a Vision-to-3D-to-Action (V-3D-A)…

机器人学 · 计算机科学 2026-05-25 Ying Chai , Litao Deng , Ruizhi Shao , Jiajun Zhang , Kangchen Lv , Liangjun Xing , Xiang Li , Hongwen Zhang , Yebin Liu

While recent Gaussian-based SLAM methods achieve photorealistic reconstruction from RGB-D data, their computational performance remains a critical bottleneck. State-of-the-art techniques operate at less than 20 fps, significantly lagging…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Zhexi Peng , Kun Zhou , Tianjia Shao