中文
相关论文

相关论文: Patch-level Gaze Distribution Prediction for Gaze …

200 篇论文

Although heatmap regression is considered a state-of-the-art method to locate facial landmarks, it suffers from huge spatial complexity and is prone to quantization error. To address this, we propose a novel attentive one-dimensional…

计算机视觉与模式识别 · 计算机科学 2020-08-28 Shi Yin , Shangfei Wang , Xiaoping Chen , Enhong Chen

Recent one-stage object detectors follow a per-pixel prediction approach that predicts both the object category scores and boundary positions from every single grid location. However, the most suitable positions for inferring different…

计算机视觉与模式识别 · 计算机科学 2022-12-02 Li Yang , Yan Xu , Shaoru Wang , Chunfeng Yuan , Ziqi Zhang , Bing Li , Weiming Hu

Many cameras implement auto-focus functionality. However, they typically require the user to manually identify the location to be focused on. While such an approach works for temporally-sparse autofocusing functionality (e.g., photo…

计算机视觉与模式识别 · 计算机科学 2017-11-10 Wolfgang Fuhl , Thiago Santini , Enkelejda Kasneci

Monocular depth predictors are typically trained on large-scale training sets which are naturally biased w.r.t the distribution of camera poses. As a result, trained predictors fail to make reliable depth predictions for testing examples…

计算机视觉与模式识别 · 计算机科学 2021-03-30 Yunhan Zhao , Shu Kong , Charless Fowlkes

Searching for small objects in large images is a task that is both challenging for current deep learning systems and important in numerous real-world applications, such as remote sensing and medical imaging. Thorough scanning of very large…

计算机视觉与模式识别 · 计算机科学 2021-04-16 Nathan Drenkow , Philippe Burlina , Neil Fendley , Onyekachi Odoemene , Jared Markowitz

Gaze prediction plays a critical role in Virtual Reality (VR) applications by reducing sensor-induced latency and enabling computationally demanding techniques such as foveated rendering, which rely on anticipating user attention. However,…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Christos Petrou , Harris Partaourides , Athanasios Balomenos , Yannis Kopsinis , Sotirios Chatzis

Enabling robots to understand human gaze target is a crucial step to allow capabilities in downstream tasks, for example, attention estimation and movement anticipation in real-world human-robot interactions. Prior works have addressed the…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Zhuangzhuang Dai , Vincent Gbouna Zakka , Luis J. Manso , Chen Li

LiDAR-based perception is critical for autonomous driving due to its robustness to poor lighting and visibility conditions. Yet, current models operate under the closed-set assumption and often fail to recognize unexpected…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Zizhao Li , Zhengkang Xiang , Jiayang Ao , Feng Liu , Joseph West , Kourosh Khoshelham

Multi-person pose estimation and tracking serve as crucial steps for video understanding. Most state-of-the-art approaches rely on first estimating poses in each frame and only then implementing data association and refinement. Despite the…

计算机视觉与模式识别 · 计算机科学 2021-06-08 Yiding Yang , Zhou Ren , Haoxiang Li , Chunluan Zhou , Xinchao Wang , Gang Hua

This paper presents a robust tracking approach to handle challenges such as occlusion and appearance change. Here, the target is partitioned into a number of patches. Then, the appearance of each patch is modeled using a dictionary composed…

计算机视觉与模式识别 · 计算机科学 2015-06-18 Ali Zarezade , Hamid R. Rabiee , Ali Soltani-Farani , Ahmad Khajenezhad

Counting objects in crowded scenes remains a challenge to computer vision. The current deep learning based approach often formulate it as a Gaussian density regression problem. Such a brute-force regression, though effective, may not…

计算机视觉与模式识别 · 计算机科学 2023-11-09 Yuehai Chen , Jing Yang , Badong Chen , Hua Gang , Shaoyi Du

Point-Level temporal action localization (PTAL) aims to localize actions in untrimmed videos with only one timestamp annotation for each action instance. Existing methods adopt the frame-level prediction paradigm to learn from the sparse…

计算机视觉与模式识别 · 计算机科学 2020-12-16 Chen Ju , Peisen Zhao , Ya Zhang , Yanfeng Wang , Qi Tian

Infrared small target detection plays an important role in the infrared search and tracking applications. In recent years, deep learning techniques were introduced to this task and achieved noteworthy effects. Following general object…

图像与视频处理 · 电气工程与系统科学 2025-07-15 Fang Chen , Chenqiang Gao , Fangcen Liu , Yue Zhao , Yuxi Zhou , Deyu Meng , Wangmeng Zuo

Memory Dependence Prediction (MDP) is a speculative technique to determine which stores, if any, a given load will depend on. Area-constrained cores are increasingly relevant in various applications such as energy-efficient or edge systems,…

Precisely detecting which object a person is paying attention to is critical for human-robot interaction since it provides important cues for the next action from the human user. We propose an end-to-end approach for gaze target detection:…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Zhi-Yi Lin , Jouh Yeong Chew , Jan van Gemert , Xucong Zhang

6D pose estimation aims at determining the object pose that best explains the camera observation. The unique solution for non-ambiguous objects can turn into a multi-modal pose distribution for symmetrical objects or when occlusions of…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Boris Meden , Asma Brazi , Fabrice Mayran de Chamisso , Steve Bourgeois , Vincent Lepetit

Gaze object prediction is a newly proposed task that aims to discover the objects being stared at by humans. It is of great application significance but still lacks a unified solution framework. An intuitive solution is to incorporate an…

计算机视觉与模式识别 · 计算机科学 2023-07-04 Binglu Wang , Tao Hu , Baoshan Li , Xiaojuan Chen , Zhijie Zhang

Gaze following aims to interpret human-scene interactions by predicting the person's focal point of gaze. Prevailing approaches often adopt a two-stage framework, whereby multi-modality information is extracted in the initial stage for gaze…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Yuehao Song , Xinggang Wang , Jingfeng Yao , Wenyu Liu , Jinglin Zhang , Xiangmin Xu

In this work, we propose a cross-view learning approach, in which images captured from a ground-level view are used as weakly supervised annotations for interpreting overhead imagery. The outcome is a convolutional neural network for…

计算机视觉与模式识别 · 计算机科学 2018-08-06 Connor Greenwell , Scott Workman , Nathan Jacobs

We present a novel multistream network that learns robust eye representations for gaze estimation. We first create a synthetic dataset containing eye region masks detailing the visible eyeball and iris using a simulator. We then perform eye…

计算机视觉与模式识别 · 计算机科学 2021-12-16 Zunayed Mahmud , Paul Hungler , Ali Etemad