English
Related papers

Related papers: SEED4D: A Synthetic Ego--Exo Dynamic 4D Data Gener…

200 papers

Current End-to-End Autonomous Driving (E2E-AD) methods resort to unifying modular designs for various tasks (e.g. perception, prediction and planning). Although optimized with a fully differentiable framework in a planning-oriented manner,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Haisheng Su , Wei Wu , Zhenjie Yang , Isabel Guan

Egocentric video provides a unique view into human perception and interaction, with growing relevance for augmented reality, robotics, and assistive technologies. However, rapid camera motion and complex scene dynamics pose major challenges…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Jan Warchocki , Xi Wang , Jonas Kulhanek , Jan van Gemert

We present EgoHumans, a new multi-view multi-human video benchmark to advance the state-of-the-art of egocentric human 3D pose estimation and tracking. Existing egocentric benchmarks either capture single subject or indoor-only scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Rawal Khirodkar , Aayush Bansal , Lingni Ma , Richard Newcombe , Minh Vo , Kris Kitani

We introduce EgoSonics, a method to generate semantically meaningful and synchronized audio tracks conditioned on silent egocentric videos. Generating audio for silent egocentric videos could open new applications in virtual reality,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Aashish Rai , Srinath Sridhar

3D scene generation has garnered growing attention in recent years and has made significant progress. Generating 4D cities is more challenging than 3D scenes due to the presence of structurally complex, visually diverse objects like…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Haozhe Xie , Zhaoxi Chen , Fangzhou Hong , Ziwei Liu

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse activities with…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Zi-Yi Dou , Xitong Yang , Tushar Nagarajan , Huiyu Wang , Jing Huang , Nanyun Peng , Kris Kitani , Fu-Jen Chu

Accurate 3D tracking of hands and their interactions with the world in unconstrained settings remains a significant challenge for egocentric computer vision. With few exceptions, existing datasets are predominantly captured in controlled…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Patrick Rim , Kun He , Kevin Harris , Braden Copple , Shangchen Han , Sizhe An , Ivan Shugurov , Tomas Hodan , He Wen , Xu Xie

In recent years, we have seen the performance of video-based person Re-Identification (ReID) methods have improved considerably. However, most of the work in this area has dealt with videos acquired by fixed cameras with wider field of…

Computer Vision and Pattern Recognition · Computer Science 2019-09-10 Emrah Basaran , Yonatan Tariku Tesfaye , Mubarak Shah

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Enrico Pallotta , Sina Mokhtarzadeh Azar , Lars Doorenbos , Serdar Ozsoy , Umar Iqbal , Juergen Gall

3D Gaussian Splatting (3DGS) has emerged as a powerful technique for real-time LiDAR and camera synthesis in autonomous driving simulation. However, simulating LiDAR with 3DGS remains challenging for extrapolated views beyond the training…

Robotics · Computer Science 2026-03-17 Yiming Huang , Xin Kang , Sipeng Zhang , Hongliang Ren , Weihua Zhang , Junjie Lai

We present SELDVisualSynth, a tool for generating synthetic videos for audio-visual sound event localization and detection (SELD). Our approach incorporates real-world background images to improve realism in synthetic audio-visual SELD data…

Sound · Computer Science 2025-04-07 Adrian S. Roman , Aiden Chang , Gerardo Meza , Iran R. Roman

Many safety-critical applications, especially in autonomous driving, require reliable object detectors. They can be very effectively assisted by a method to search for and identify potential failures and systematic errors before these…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Valentyn Boreiko , Matthias Hein , Jan Hendrik Metzen

Cross-view video synthesis task seeks to generate video sequences of one view from another dramatically different view. In this paper, we investigate the exocentric (third-person) view to egocentric (first-person) view video generation…

Computer Vision and Pattern Recognition · Computer Science 2021-07-08 Gaowen Liu , Hao Tang , Hugo Latapie , Jason Corso , Yan Yan

The accelerating development of autonomous driving technology has placed greater demands on obtaining large amounts of high-quality data. Representative, labeled, real world data serves as the fuel for training deep learning networks,…

Computer Vision and Pattern Recognition · Computer Science 2021-12-24 Pengchuan Xiao , Zhenlei Shao , Steven Hao , Zishuo Zhang , Xiaolin Chai , Judy Jiao , Zesong Li , Jian Wu , Kai Sun , Kun Jiang , Yunlong Wang , Diange Yang

Egocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection…

In this paper, we propose a novel learning approach for feed-forward one-shot 4D head avatar synthesis. Different from existing methods that often learn from reconstructing monocular videos guided by 3DMM, we employ pseudo multi-view videos…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Yu Deng , Duomin Wang , Baoyuan Wang

Diffusion models have recently enabled precise and photorealistic facial editing across a wide range of semantic attributes. Beyond single-step modifications, a growing class of applications now demands the ability to analyze and track…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Yule Zhu , Ping Liu , Zhedong Zheng , Wei Liu

Once an academic venture, autonomous driving has received unparalleled corporate funding in the last decade. Still, the operating conditions of current autonomous cars are mostly restricted to ideal scenarios. This means that driving in…

Computer Vision and Pattern Recognition · Computer Science 2021-03-11 Mathias Gehrig , Willem Aarents , Daniel Gehrig , Davide Scaramuzza

Generating dynamic 3D object from a single-view video is challenging due to the lack of 4D labeled data. An intuitive approach is to extend previous image-to-3D pipelines by transferring off-the-shelf image generation models such as score…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Zijie Pan , Zeyu Yang , Xiatian Zhu , Li Zhang

This article presents a synthetic distracted driving (SynDD2 - a continuum of SynDD1) dataset for machine learning models to detect and analyze drivers' various distracted behavior and different gaze zones. We collected the data in a…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Mohammed Shaiqur Rahman , Jiyang Wang , Senem Velipasalar Gursoy , David Anastasiu , Shuo Wang , Anuj Sharma