English
Related papers

Related papers: G2D: from GTA to Data

200 papers

Temporal action detection is a fundamental yet challenging task in video understanding. Video context is a critical cue to effectively detect actions, but current works mainly focus on temporal context, while neglecting semantic context as…

Computer Vision and Pattern Recognition · Computer Science 2020-04-06 Mengmeng Xu , Chen Zhao , David S. Rojas , Ali Thabet , Bernard Ghanem

End-to-end autonomous driving has evolved from the conventional paradigm based on sparse perception into vision-language-action (VLA) models, which focus on learning language descriptions as an auxiliary task to facilitate planning. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Sicheng Zuo , Zixun Xie , Wenzhao Zheng , Shaoqing Xu , Fang Li , Hanbing Li , Long Chen , Zhi-Xin Yang , Jiwen Lu

We present a real-time approach for image-based localization within large scenes that have been reconstructed offline using structure from motion (Sfm). From monocular video, our method continuously computes a precise 6-DOF camera pose, by…

Computer Vision and Pattern Recognition · Computer Science 2015-03-20 Hyon Lim , Sudipta Sinha , Michael Cohen , Matt Uyttendaele

State-of-the-art novel view synthesis methods achieve impressive results for multi-view captures of static 3D scenes. However, the reconstructed scenes still lack "liveliness," a key component for creating engaging 3D experiences. Recently,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Thomas Wimmer , Michael Oechsle , Michael Niemeyer , Federico Tombari

This paper presents GSWorld, a robust, photo-realistic simulator for robotics manipulation that combines 3D Gaussian Splatting with physics engines. Our framework advocates "closing the loop" of developing manipulation policies with…

Robotics · Computer Science 2025-10-24 Guangqi Jiang , Haoran Chang , Ri-Zhao Qiu , Yutong Liang , Mazeyu Ji , Jiyue Zhu , Zhao Dong , Xueyan Zou , Xiaolong Wang

Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based methods create 3D scenes from driving logs and synthesize…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Jiageng Mao , Boyi Li , Boris Ivanovic , Yuxiao Chen , Yan Wang , Yurong You , Chaowei Xiao , Danfei Xu , Marco Pavone , Yue Wang

For driver observation frameworks, clean datasets collected in controlled simulated environments often serve as the initial training ground. Yet, when deployed under real driving conditions, such simulator-trained models quickly face the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Walter Morales-Alvarez , Novel Certad , Alina Roitberg , Rainer Stiefelhagen , Cristina Olaverri-Monreal

Urban scene reconstruction requires modeling both static infrastructure and dynamic elements while supporting diverse environmental conditions. We present \textbf{StyledStreets}, a multi-style street simulator that achieves…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Yuyin Chen , Yida Wang , Xueyang Zhang , Kun Zhan , Peng Jia , Yifei Zhan , Xianpeng Lang

Autonomous driving needs fast, scalable 4D reconstruction and re-simulation for training and evaluation, yet most methods for dynamic driving scenes still rely on per-scene optimization, known camera calibration, or short frame windows,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Xiaoxue Chen , Ziyi Xiong , Yuantao Chen , Gen Li , Nan Wang , Hongcheng Luo , Long Chen , Haiyang Sun , Bing Wang , Guang Chen , Hangjun Ye , Hongyang Li , Ya-Qin Zhang , Hao Zhao

The aim of this paper is to propose new algorithms for Field of Vision (FOV) computation which improve on existing work at high resolutions. FOV refers to the set of locations that are visible from a specific position in a scene of a…

Computer Vision and Pattern Recognition · Computer Science 2021-01-28 Evan R. M. Debenham , Roberto Solis-Oba

Recent high-performing image-to-video (I2V) models based on variants of the diffusion transformer (DiT) have displayed remarkable inherent world-modeling capabilities by virtue of training on large scale video datasets. We investigate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Aaron Appelle , Jerome P. Lynch

Building recognition and 3D reconstruction of human made structures in urban scenarios has become an interesting and actual topic in the image processing domain. For this research topic the Computer Vision and Augmented Reality areas…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Orhei Ciprian , Vert Silviu , Mocofan Muguras , Vasiu Radu

Digital twin is a problem of augmenting real objects with their digital counterparts. It can underpin a wide range of applications in augmented reality (AR), autonomy, and UI/UX. A critical component in a good digital-twin system is…

Computer Vision and Pattern Recognition · Computer Science 2023-04-13 Weiyu Feng , Seth Z. Zhao , Chuanyu Pan , Adam Chang , Yichen Chen , Zekun Wang , Allen Y. Yang

Traffic congestion in urban areas presents significant challenges, and Intelligent Transportation Systems (ITS) have sought to address these via automated and adaptive controls. However, these systems often struggle to transfer simulated…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Daniel Rodriguez-Criado , Maria Chli , Luis J. Manso , George Vogiatzis

Collaborative driving systems leverage vehicle-to-everything (V2X) communication for multi-agent collaborative perception to enhance driving safety, yet they remain constrained by scarce annotated real-world V2X driving datasets and limited…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yihang Tao , Yu Guo , Senkang Hu , Yanan Ma , Zihan Fang , Sam Kwong , Yuguang Fang

Different video understanding tasks are typically treated in isolation, and even with distinct types of curated data (e.g., classifying sports in one dataset, tracking animals in another). However, in wearable cameras, the immersive…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

Embodied intelligence requires high-fidelity simulation environments to support perception and decision-making, yet existing platforms often suffer from data contamination and limited flexibility. To mitigate this, we propose…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Lechao Zhang , Haoran Xu , Jingyu Gong , Xuhong Wang , Yuan Xie , Xin Tan

For 6-DoF grasp detection, simulated data is expandable to train more powerful model, but it faces the challenge of the large gap between simulation and real world. Previous works bridge this gap with a sim-to-real way. However, this way…

Robotics · Computer Science 2024-10-10 Jia-Feng Cai , Zibo Chen , Xiao-Ming Wu , Jian-Jian Jiang , Yi-Lin Wei , Wei-Shi Zheng

Synthesizing controllable 6-DOF object manipulation trajectories in 3D environments is essential for enabling robots to interact with complex scenes, yet remains challenging due to the need for accurate spatial reasoning, physical…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Huajian Zeng , Abhishek Saroha , Daniel Cremers , Xi Wang

Generating complex multi-actor scenario videos remains difficult even for state-of-the-art neural generators, while evaluating them is hard due to the lack of ground truth for physical plausibility and semantic faithfulness. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Nicolae Cudlenco , Mihai Masala , Marius Leordeanu