English
Related papers

Related papers: EPIC Fields: Marrying 3D Geometry and Video Unders…

200 papers

The pipeline of current robotic pick-and-place methods typically consists of several stages: grasp pose detection, finding inverse kinematic solutions for the detected poses, planning a collision-free trajectory, and then executing the…

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li

In this report, we describe the technical details of our submission for the EPIC-Kitchen-100 action anticipation challenge. Our modelings, the higher-order recurrent space-time transformer and the message-passing neural network with edge…

Computer Vision and Pattern Recognition · Computer Science 2022-06-23 Tsung-Ming Tai , Oswald Lanz , Giuseppe Fiameni , Yi-Kwan Wong , Sze-Sen Poon , Cheng-Kuang Lee , Ka-Chun Cheung , Simon See

Neural representations have emerged as a new paradigm for applications in rendering, imaging, geometric modeling, and simulation. Compared to traditional representations such as meshes, point clouds, or volumes they can be flexibly…

Computer Vision and Pattern Recognition · Computer Science 2021-05-07 Julien N. P. Martel , David B. Lindell , Connor Z. Lin , Eric R. Chan , Marco Monteiro , Gordon Wetzstein

Lifting perspective images and videos to 360{\deg} panoramas enables immersive 3D world generation. Existing approaches often rely on explicit geometric alignment between the perspective and the equirectangular projection (ERP) space. Yet,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Ziyi Wu , Daniel Watson , Andrea Tagliasacchi , David J. Fleet , Marcus A. Brubaker , Saurabh Saxena

Deep learning has enabled remarkable improvements in grasp synthesis for previously unseen objects from partial object views. However, existing approaches lack the ability to explicitly reason about the full 3D geometry of the object when…

Robotics · Computer Science 2020-03-19 Mark Van der Merwe , Qingkai Lu , Balakumar Sundaralingam , Martin Matak , Tucker Hermans

Segment matching is an important intermediate task in computer vision that establishes correspondences between semantically or geometrically coherent regions across images. Unlike keypoint matching, which focuses on localized features,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Rohit Jayanti , Swayam Agrawal , Vansh Garg , Siddharth Tourani , Muhammad Haris Khan , Sourav Garg , Madhava Krishna

Current 3D scene understanding methods are limited by offline-collected multi-view data or pre-constructed 3D geometry. In this paper, we present ExtractAnything3D (EA3D), a unified online framework for open-world 3D object extraction that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Xiaoyu Zhou , Jingqi Wang , Yuang Jia , Yongtao Wang , Deqing Sun , Ming-Hsuan Yang

We present a novel single-stage framework, Neural Photon Field (NePF), to address the ill-posed inverse rendering from multi-view images. Contrary to previous methods that recover the geometry, material, and illumination in multiple stages…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Tuen-Yue Tsui , Qin Zou

Compared to frame-based methods, computational neuromorphic imaging using event cameras offers significant advantages, such as minimal motion blur, enhanced temporal resolution, and high dynamic range. The multi-view consistency of Neural…

Computer Vision and Pattern Recognition · Computer Science 2025-01-08 Chaoran Feng , Wangbo Yu , Xinhua Cheng , Zhenyu Tang , Junwu Zhang , Li Yuan , Yonghong Tian

We focus on multi-modal fusion for egocentric action recognition, and propose a novel architecture for multi-modal temporal-binding, i.e. the combination of modalities within a range of temporal offsets. We train the architecture with three…

Computer Vision and Pattern Recognition · Computer Science 2019-08-23 Evangelos Kazakos , Arsha Nagrani , Andrew Zisserman , Dima Damen

Since the advent of Neural Radiance Fields, novel view synthesis has received tremendous attention. The existing approach for the generalization of radiance field reconstruction primarily constructs an encoding volume from nearby source…

Computer Vision and Pattern Recognition · Computer Science 2023-08-09 Jingliang Li , Qiang Zhou , Chaohui Yu , Zhengda Lu , Jun Xiao , Zhibin Wang , Fan Wang

Event cameras are rapidly emerging as powerful vision sensors for 3D reconstruction, uniquely capable of asynchronously capturing per-pixel brightness changes. Compared to traditional frame-based cameras, event cameras produce sparse yet…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Chuanzhi Xu , Haoxian Zhou , Langyi Chen , Haodong Chen , Zeke Zexi Hu , Zhicheng Lu , Ying Zhou , Vera Chung , Qiang Qu , Weidong Cai

Human activities are inherently complex, often involving numerous object interactions. To better understand these activities, it is crucial to model their interactions with the environment captured through dynamic changes. The recent…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Daiwei Zhang , Gengyan Li , Jiajie Li , Mickaël Bressieux , Otmar Hilliges , Marc Pollefeys , Luc Van Gool , Xi Wang

We introduce the first learning-based dense matching algorithm, termed Equirectangular Projection-Oriented Dense Kernelized Feature Matching (EDM), specifically designed for omnidirectional images. Equirectangular projection (ERP) images,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Dongki Jung , Jaehoon Choi , Yonghan Lee , Somi Jeong , Taejae Lee , Dinesh Manocha , Suyong Yeon

Recent advances in video analytics address real-time data drift by continuously retraining specialized, lightweight DNN models for individual cameras. However, the current practice of retraining a separate model for each camera suffers from…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-15 Yuze He , Ferdi Kossmann , Srinivasan Seshan , Peter Steenkiste

The filming of sporting events projects and flattens the movement of athletes in the world onto a 2D broadcast image. The pixel locations of joints in these images can be detected with high validity. Recovering the actual 3D movement of the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Tobias Baumgartner , Stefanie Klatt

Neural reconstruction models for autonomous driving simulation have made significant strides in recent years, with dynamic models becoming increasingly prevalent. However, these models are typically limited to handling in-domain objects…

For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images. We address this challenge by introducing Hypersim, a photorealistic synthetic dataset for holistic…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Mike Roberts , Jason Ramapuram , Anurag Ranjan , Atulit Kumar , Miguel Angel Bautista , Nathan Paczan , Russ Webb , Joshua M. Susskind

Neural rendering is a new image and video generation method based on deep learning. It combines the deep learning model with the physical knowledge of computer graphics, to obtain a controllable and realistic scene model, and realize the…

Graphics · Computer Science 2024-02-02 Xinkai Yan , Jieting Xu , Yuchi Huo , Hujun Bao