中文
相关论文

相关论文: EPIC Fields: Marrying 3D Geometry and Video Unders…

200 篇论文

A spike camera is a specialized high-speed visual sensor that offers advantages such as high temporal resolution and high dynamic range compared to conventional frame cameras. These features provide the camera with significant advantages in…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Jinze Yu , Xin Peng , Zhengda Lu , Laurent Kneip , Yiqun Wang

Neural fields have emerged as a powerful paradigm for representing various signals, including videos. However, research on improving the parameter efficiency of neural fields is still in its early stages. Even though neural fields that map…

计算机视觉与模式识别 · 计算机科学 2022-10-06 Daniel Rho , Junwoo Cho , Jong Hwan Ko , Eunbyung Park

Computer vision is largely based on 2D techniques, with 3D vision still relegated to a relatively narrow subset of applications. However, by building on recent advances in 3D models such as neural radiance fields, some authors have shown…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Vadim Tschernezki , Diane Larlus , Iro Laina , Andrea Vedaldi

In this report, we present the Baidu-UTS submission to the EPIC-Kitchens Action Recognition Challenge in CVPR 2019. This is the winning solution to this challenge. In this task, the goal is to predict verbs, nouns, and actions from the…

计算机视觉与模式识别 · 计算机科学 2019-06-25 Xiaohan Wang , Yu Wu , Linchao Zhu , Yi Yang

In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge. Building upon the video-language two-tower model and leveraging…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Baoqi Pei , Guo Chen , Jilan Xu , Yuping He , Yicheng Liu , Kanghua Pan , Yifei Huang , Yali Wang , Tong Lu , Limin Wang , Yu Qiao

Action recognition is currently one of the top-challenging research fields in computer vision. Convolutional Neural Networks (CNNs) have significantly boosted its performance but rely on fixed-size spatio-temporal windows of analysis,…

计算机视觉与模式识别 · 计算机科学 2020-08-27 Alejandro López-Cifuentes , Marcos Escudero-Viñolo , Jesús Bescós

Spatiotemporal video grounding aims to localize target entities in videos based on textual queries. While existing research has made significant progress in exocentric videos, the egocentric setting remains relatively underexplored, despite…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Shuo Liang , Yiwu Zhong , Zi-Yuan Hu , Yeyao Tao , Liwei Wang

While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on question-answering or multiple-choice formats. These protocols allow models to exploit…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Haozhe Shan , Xiancong Ren , Han Dong , Haoyuan Shi , Yingji Zhang , Jiayu Hu , Yi Zhang , Yong Dai , Bin Shen , Lizhen Qu , Zenglin Xu , Xiaozhu Ju

Egocentric interaction perception is one of the essential branches in investigating human-environment interaction, which lays the basis for developing next-generation intelligent systems. However, existing egocentric interaction…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Yuejiao Su , Yi Wang , Qiongyang Hu , Chuang Yang , Lap-Pui Chau

Neural radiance fields (NeRF) appeared recently as a powerful tool to generate realistic views of objects and confined areas. Still, they face serious challenges with open scenes, where the camera has unrestricted movement and content can…

计算机视觉与模式识别 · 计算机科学 2023-03-23 Ahmad AlMughrabi , Umair Haroon , Ricardo Marques , Petia Radeva

Video generation models have progressed tremendously through large latent diffusion transformers trained with rectified flow techniques. Yet these models still struggle with geometric inconsistencies, unstable motion, and visual artifacts…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Orest Kupyn , Fabian Manhardt , Federico Tombari , Christian Rupprecht

Spatial intelligence, encompassing 3D reconstruction, perception, and reasoning, is fundamental to applications such as robotics, aerial imaging, and extended reality. A key enabler is the real-time, accurate estimation of core 3D…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Wenyan Cong , Yiqing Liang , Yancheng Zhang , Ziyi Yang , Yan Wang , Boris Ivanovic , Marco Pavone , Chen Chen , Zhangyang Wang , Zhiwen Fan

A new unified video analytics framework (ER3) is proposed for complex event retrieval, recognition and recounting, based on the proposed video imprint representation, which exploits temporal correlations among image features across video…

计算机视觉与模式识别 · 计算机科学 2021-06-08 Zhanning Gao , Le Wang , Nebojsa Jojic , Zhenxing Niu , Nanning Zheng , Gang Hua

Recent advances in machine learning have created increasing interest in solving visual computing problems using a class of coordinate-based neural networks that parametrize physical properties of scenes or objects across space and time.…

计算机视觉与模式识别 · 计算机科学 2022-04-07 Yiheng Xie , Towaki Takikawa , Shunsuke Saito , Or Litany , Shiqin Yan , Numair Khan , Federico Tombari , James Tompkin , Vincent Sitzmann , Srinath Sridhar

We present Neural Feature Fusion Fields (N3F), a method that improves dense 2D image feature extractors when the latter are applied to the analysis of multiple images reconstructible as a 3D scene. Given an image feature extractor, for…

计算机视觉与模式识别 · 计算机科学 2022-09-09 Vadim Tschernezki , Iro Laina , Diane Larlus , Andrea Vedaldi

Photo-realistic rendering and novel view synthesis play a crucial role in human-computer interaction tasks, from gaming to path planning. Neural Radiance Fields (NeRFs) model scenes as continuous volumetric functions and achieve remarkable…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Iryna Repinetska , Anna Hilsmann , Peter Eisert

Accurate 3D geometric perception is an important prerequisite for a wide range of spatial AI systems. While state-of-the-art methods depend on large-scale training data, acquiring consistent and precise 3D annotations from in-the-wild…

Reconstruction of deformable scenes from endoscopic videos is important for many applications such as intraoperative navigation, surgical visual perception, and robotic surgery. It is a foundational requirement for realizing autonomous…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Shreya Saha , Zekai Liang , Shan Lin , Jingpei Lu , Michael Yip , Sainan Liu

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view…

Neural volumetric representations such as Neural Radiance Fields (NeRF) have emerged as a compelling technique for learning to represent 3D scenes from images with the goal of rendering photorealistic images of the scene from unobserved…

计算机视觉与模式识别 · 计算机科学 2021-03-29 Peter Hedman , Pratul P. Srinivasan , Ben Mildenhall , Jonathan T. Barron , Paul Debevec