English
Related papers

Related papers: GS-Playground: A High-Throughput Photorealistic Si…

200 papers

The remote sensing image intelligence understanding model is undergoing a new profound paradigm shift which has been promoted by multi-modal large language model (MLLM), i.e. from the paradigm learning a domain model (LaDM) shifts to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Linrui Xu , Ling Zhao , Wang Guo , Qiujun Li , Kewang Long , Kaiqi Zou , Yuhan Wang , Haifeng Li

In-the-wild photo collections often contain limited volumes of imagery and exhibit multiple appearances, e.g., taken at different times of day or seasons, posing significant challenges to scene reconstruction and novel view synthesis.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Deming Li , Kaiwen Jiang , Yutao Tang , Ravi Ramamoorthi , Rama Chellappa , Cheng Peng

Robots rely heavily on sensors, especially RGB and depth cameras, to perceive and interact with the world. RGB cameras record 2D images with rich semantic information while missing precise spatial information. On the other side, depth…

Robotics · Computer Science 2023-10-16 Tong Zhang , Yingdong Hu , Hanchen Cui , Hang Zhao , Yang Gao

Despite rapid progress, embodied agents still struggle with long-horizon manipulation that requires maintaining spatial consistency, causal dependencies, and goal constraints. A key limitation of existing approaches is that task reasoning…

Physical interactive robotics, ranging from wearable devices to collaborative humanoid robots, require close coordination between mechanical design and control. However, evaluating interactive dynamics is challenging due to complex human…

Robotics · Computer Science 2026-03-11 Chenhui Zuo , Jinhao Xu , Michael Qian Vergnolle , Yanan Sui

Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms. Existing cross-embodiment approaches try to map human motion to robot actions, but often…

This paper investigates the task of the open-ended interactive robotic manipulation on table-top scenarios. While recent Large Language Models (LLMs) enhance robots' comprehension of user instructions, their lack of visual grounding…

Robotics · Computer Science 2024-08-16 Tianyu Wang , Haitao Lin , Junqiu Yu , Yanwei Fu

Developing visual perception models for active agents and sensorimotor control are cumbersome to be done in the physical world, as existing algorithms are too slow to efficiently learn in real-time and robots are fragile and costly. This…

Artificial Intelligence · Computer Science 2018-09-03 Fei Xia , Amir Zamir , Zhi-Yang He , Alexander Sax , Jitendra Malik , Silvio Savarese

Semantic-aware 3D scene reconstruction is essential for autonomous robots to perform complex interactions. Semantic SLAM, an online approach, integrates pose tracking, geometric reconstruction, and semantic mapping into a unified framework,…

Robotics · Computer Science 2025-05-20 Zuxing Lu , Xin Yuan , Shaowen Yang , Jingyu Liu , Changyin Sun

Teaching robots dexterous manipulation skills often requires collecting hundreds of demonstrations using wearables or teleoperation, a process that is challenging to scale. Videos of human-object interactions are easier to collect and…

Robotics · Computer Science 2025-08-19 Tyler Ga Wei Lum , Olivia Y. Lee , C. Karen Liu , Jeannette Bohg

Recent progress in legged locomotion has allowed highly dynamic and parkour-like behaviors for robots, similar to their biological counterparts. Yet, these methods mostly rely on egocentric (first-person) perception, limiting their…

Robotics · Computer Science 2025-12-01 Rémy Rahem , Wael Suleiman

Embodied intelligence requires precise reconstruction and rendering to simulate large-scale real-world data. Although 3D Gaussian Splatting (3DGS) has recently demonstrated high-quality results with real-time performance, it still faces…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Haodong Xiang , Xinghui Li , Kai Cheng , Xiansong Lai , Wanting Zhang , Zhichao Liao , Long Zeng , Xueping Liu

Machine learning has facilitated significant advancements across various robotics domains, including navigation, locomotion, and manipulation. Many such achievements have been driven by the extensive use of simulation as a critical tool for…

3D Gaussian splatting (3D-GS) has recently revolutionized novel view synthesis in the simultaneous localization and mapping (SLAM) problem. However, most existing algorithms fail to fully capture the underlying structure, resulting in…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Tianci Wen , Zhiang Liu , Yongchun Fang

In this work, we investigate an important task named instruction-following text embedding, which generates dynamic text embeddings that adapt to user instructions, highlighting specific attributes of text. Despite recent advancements,…

Computation and Language · Computer Science 2025-06-02 Yingchaojie Feng , Yiqun Sun , Yandong Sun , Minfeng Zhu , Qiang Huang , Anthony K. H. Tung , Wei Chen

Neural scene representations such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have transformed how 3D environments are modeled, rendered, and interpreted. NeRF introduced view-consistent photorealism via volumetric…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Javed Ahmad , Penggang Gao , Donatien Delehelle , Mennuti Canio , Nikhil Deshpande , Jesús Ortiz , Darwin G. Caldwell , Yonas Teodros Tefera

Recently, the 3D Gaussian splatting (3DGS) technique for real-time radiance field rendering has revolutionized the field of volumetric scene representation, providing users with an immersive experience. But in return, it also poses a large…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Zhiye Tang , Qiudan Zhang , Lei Zhang , Junhui Hou , You Yang , Xu Wang

Recently, the multi-modal fusion of RGB, depth, and semantics has shown great potential in dense Simultaneous Localization and Mapping (SLAM). However, a prerequisite for generating consistent semantic maps is the availability of dense,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Linfei Li , Lin Zhang , Zhong Wang , Ying Shen

Neural simulators promise efficient surrogates for physics simulation, but scaling them is bottlenecked by the prohibitive cost of generating high-fidelity training data. Pre-training on abundant off-the-shelf geometries offers a natural…

Machine Learning · Computer Science 2026-05-21 Haixu Wu , Minghao Guo , Zongyi Li , Zhiyang Dou , Mingsheng Long , Kaiming He , Wojciech Matusik

We propose GS-IR, a novel inverse rendering approach based on 3D Gaussian Splatting (GS) that leverages forward mapping volume rendering to achieve photorealistic novel view synthesis and relighting results. Unlike previous works that use…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Zhihao Liang , Qi Zhang , Ying Feng , Ying Shan , Kui Jia
‹ Prev 1 4 5 6 7 8 10 Next ›