中文
相关论文

相关论文: GazePrior: Zero-Shot AR/VR Eye Tracking via Learne…

200 篇论文

The rapid growth of stereoscopic displays, including VR headsets and 3D cinemas, has led to increasing demand for high-quality stereo video content. However, producing 3D videos remains costly and complex, while automatic…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Guibao Shen , Yihua Du , Wenhang Ge , Jing He , Chirui Chang , Donghao Zhou , Zhen Yang , Luozhou Wang , Xin Tao , Ying-Cong Chen

We introduce G-CUT3R, a novel feed-forward approach for guided 3D scene reconstruction that enhances the CUT3R model by integrating prior information. Unlike existing feed-forward methods that rely solely on input images, our method…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ramil Khafizov , Artem Komarichev , Ruslan Rakhimov , Peter Wonka , Evgeny Burnaev

This work presents a next-generation human-robot interface that can infer and realize the user's manipulation intention via sight only. Specifically, we develop a system that integrates near-eye-tracking and robotic manipulation to enable…

机器人学 · 计算机科学 2023-05-16 Shaochen Wang , Wei Zhang , Zhangli Zhou , Jiaxi Cao , Ziyang Chen , Kang Chen , Bin Li , Zhen Kan

We propose GazeNeRF, a 3D-aware method for the task of gaze redirection. Existing gaze redirection methods operate on 2D images and struggle to generate 3D consistent results. Instead, we build on the intuition that the face region and…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Alessandro Ruzzi , Xiangwei Shi , Xi Wang , Gengyan Li , Shalini De Mello , Hyung Jin Chang , Xucong Zhang , Otmar Hilliges

Appearance-based supervised methods with full-face image input have made tremendous advances in recent gaze estimation tasks. However, intensive human annotation requirement inhibits current methods from achieving industrial level accuracy…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Yangzhou Jiang , Yinxin Lin , Yaoming Wang , Teng Li , Bilian Ke , Bingbing Ni

Feed-forward foundation models for multi-view 3-dimensional (3D) reconstruction have been trained on large-scale datasets of perspective images; when tested on wide field-of-view images, e.g., from a fisheye camera, their performance…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Ruxiao Duan , Erin Hong , Dongxu Zhao , Eric Turner , Alex Wong , Yunwen Zhou

Gaze communication plays a crucial role in daily social interactions. Quantifying this behavior can help in human-computer interaction and digital phenotyping. While end-to-end models exist for gaze target detection, they only utilize a…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Ryan Anthony Jalova de Belen , Gelareh Mohammadi , Arcot Sowmya

The performance of existing single-view 3D reconstruction methods heavily relies on large-scale 3D annotations. However, such annotations are tedious and expensive to collect. Semi-supervised learning serves as an alternative way to…

计算机视觉与模式识别 · 计算机科学 2022-10-03 Zhen Xing , Hengduo Li , Zuxuan Wu , Yu-Gang Jiang

Virtual reality (VR) is not a new technology but has been in development for decades, driven by advances in computer technology. Currently, VR technology is increasingly being used in applications to enable immersive, yet controlled…

人机交互 · 计算机科学 2022-11-24 Hong Gao

In this paper, we present a survey of deep learning-based methods for the regression of gaze direction vector from head and eye images. We describe in detail numerous published methods with a focus on the input data, architecture of the…

计算机视觉与模式识别 · 计算机科学 2024-10-23 Franko Šikić , Donik Vršnak , Sven Lončarić

Multi-view aggregation promises to overcome the occlusion and missed detection challenge in multi-object detection and tracking. Recent approaches in multi-view detection and 3D object detection made a huge performance leap by projecting…

计算机视觉与模式识别 · 计算机科学 2023-10-23 Torben Teepe , Philipp Wolters , Johannes Gilg , Fabian Herzog , Gerhard Rigoll

We present UnrealEgo, i.e., a new large-scale naturalistic dataset for egocentric 3D human pose estimation. UnrealEgo is based on an advanced concept of eyeglasses equipped with two fisheye cameras that can be used in unconstrained…

计算机视觉与模式识别 · 计算机科学 2022-08-03 Hiroyasu Akada , Jian Wang , Soshi Shimada , Masaki Takahashi , Christian Theobalt , Vladislav Golyanik

We propose the notion of Attention-Aware Visualizations (AAVs) that track the user's perception of a visual representation over time and feed this information back to the visualization. Such context awareness is particularly useful for…

Understanding human behavior is key for robots and intelligent systems that share a space with people. Accordingly, research that enables such systems to perceive, track, learn and predict human behavior as well as to plan and interact with…

Vision foundation models have attracted significant attention for their ability to leverage large-scale unlabeled visual data. This advantage is particularly important in remote sensing, where data acquisition is costly and annotation often…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Hyobin Park , Minseok Seo , Dong-Geol Choi

Deep robot vision models are widely used for recognizing objects from camera images, but shows poor performance when detecting objects at untrained positions. Although such problem can be alleviated by training with large datasets, the…

机器人学 · 计算机科学 2022-10-26 Hyogo Hiruma , Hiroki Mori , Hiroshi Ito , Tetsuya Ogata

Estimating the 6D pose of novel objects is a fundamental yet challenging problem in robotics, often relying on access to object CAD models. However, acquiring such models can be costly and impractical. Recent approaches aim to bypass this…

机器人学 · 计算机科学 2025-08-25 Zhaodong Jiang , Ashish Sinha , Tongtong Cao , Yuan Ren , Bingbing Liu , Binbin Xu

Recent 3D generation models typically rely on limited-scale 3D `gold-labels' or 2D diffusion priors for 3D content creation. However, their performance is upper-bounded by constrained 3D priors due to the lack of scalable learning…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Baorui Ma , Huachen Gao , Haoge Deng , Zhengxiong Luo , Tiejun Huang , Lulu Tang , Xinlong Wang

Generating a bird's eye view of road users is beneficial for a variety of applications, including navigation, detecting agent conflicts, and measuring space occupancy, as well as the ability to utilise the metric system to measure distances…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Mohamed R Ibrahim

Large language and vision models have been leading a revolution in visual computing. By greatly scaling up sizes of data and model parameters, the large models learn deep priors which lead to remarkable performance in various tasks. In this…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Junsheng Zhou , Yu-Shen Liu , Zhizhong Han