English
Related papers

Related papers: End-to-End Human-Gaze-Target Detection with Transf…

200 papers

Human gaze provides valuable information on human focus and intentions, making it a crucial area of research. Recently, deep learning has revolutionized appearance-based gaze estimation. However, due to the unique features of gaze…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Yihua Cheng , Haofei Wang , Yiwei Bao , Feng Lu

Learning-based gaze estimation methods require large amounts of training data with accurate gaze annotations. Facing such demanding requirements of gaze data collection and annotation, several image synthesis methods were proposed, which…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Shiwei Jin , Zhen Wang , Lei Wang , Ning Bi , Truong Nguyen

Object detection in videos plays a crucial role in advancing applications such as public safety and anomaly detection. Existing methods have explored different techniques, including CNN, deep learning, and Transformers, for object detection…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Badhan Chandra Das , M. Hadi Amini , Yanzhao Wu

Recent DEtection TRansformer (DETR) based frameworks have achieved remarkable success in end-to-end object detection. However, the reliance on the Hungarian algorithm for bipartite matching between queries and ground truths introduces…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Shoumeng Qiu , Xinrun Li , Yang Long

Gait recognition is a remote biometric technology that utilizes the dynamic characteristics of human movement to identify individuals even under various extreme lighting conditions. Due to the limitation in spatial perception capability…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Jiaxing Hao , Yanxi Wang , Zhigang Chang , Hongmin Gao , Zihao Cheng , Chen Wu , Xin Zhao , Peiye Fang , Rachmat Muwardi

Active perception and foveal vision are the foundations of the human visual system. While foveal vision reduces the amount of information to process during a gaze fixation, active perception will change the gaze direction to the most…

Computer Vision and Pattern Recognition · Computer Science 2022-08-25 Alexandre M. F. Dias , Luís Simões , Plinio Moreno , Alexandre Bernardino

Video salient object detection aims at discovering the most visually distinctive objects in a video. How to effectively take object motion into consideration during video salient object detection is a critical issue. Existing…

Computer Vision and Pattern Recognition · Computer Science 2019-10-04 Haofeng Li , Guanqi Chen , Guanbin Li , Yizhou Yu

Edge detection has long been an important problem in the field of computer vision. Previous works have explored category-agnostic or category-aware edge detection. In this paper, we explore edge detection in the context of object instances.…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Xueyan Zou , Haotian Liu , Yong Jae Lee

Deep neural networks have demonstrated superior performance on appearance-based gaze estimation tasks. However, due to variations in person, illuminations, and background, performance degrades dramatically when applying the model to a new…

Computer Vision and Pattern Recognition · Computer Science 2022-10-06 Ruicong Liu , Yiwei Bao , Mingjie Xu , Haofei Wang , Yunfei Liu , Feng Lu

Gaze and face tracking algorithms have traditionally battled a compromise between computational complexity and accuracy; the most accurate neural net algorithms cannot be implemented in real time, but less complex real-time algorithms…

Computer Vision and Pattern Recognition · Computer Science 2017-11-21 George He , Sami Oueida , Tucker Ward

Grasp detection methods typically target the detection of a set of free-floating hand poses that can grasp the object. However, not all of the detected grasp poses are executable due to physical constraints. Even though it is…

Robotics · Computer Science 2025-08-06 Tianyi Ko , Takuya Ikeda , Balazs Opra , Koichi Nishiwaki

The High-Resolution Transformer (HRFormer) can maintain high-resolution representation and share global receptive fields. It is friendly towards salient object detection (SOD) in which the input and output have the same resolution. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-01-25 Bin Tang , Zhengyi Liu , Yacheng Tan , Qian He

Visual Geometry Grounded Transformer (VGGT) has advanced 3D vision, yet its global attention layers suffer from quadratic computational costs that hinder scalability. Several sparsification-based acceleration techniques have been proposed…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yongsung Kim , Wooseok Song , Jaihyun Lew , Hun Hwangbo , Jaehoon Lee , Sungroh Yoon

Transformers have proven superior performance for a wide variety of tasks since they were introduced. In recent years, they have drawn attention from the vision community in tasks such as image classification and object detection. Despite…

Computer Vision and Pattern Recognition · Computer Science 2022-10-03 Yihong Xu , Yutong Ban , Guillaume Delorme , Chuang Gan , Daniela Rus , Xavier Alameda-Pineda

A user's eyes provide means for Human Computer Interaction (HCI) research as an important modal. The time to time scientific explorations of the eye has already seen an upsurge of the benefits in HCI applications from gaze estimation to the…

Computer Vision and Pattern Recognition · Computer Science 2021-06-10 Atul Sahay , Imon Mukherjee , Kavi Arya

Gaze object prediction (GOP) aims to predict the category and location of the object that a human is looking at. Previous methods utilized box-level supervision to identify the object that a person is looking at, but struggled with semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Yang Jin , Lei Zhang , Shi Yan , Bin Fan , Binglu Wang

In this paper, we present the first transformer-based model to address the challenging problem of egocentric gaze estimation. We observe that the connection between the global scene context and local visual information is vital for…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Bolin Lai , Miao Liu , Fiona Ryan , James M. Rehg

In the technical report, we present a novel transformer-based framework for nuScenes lidar-based object detection task, termed Spatial Expansion Group Transformer (SEGT). To efficiently handle the irregular and sparse nature of point cloud,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Cheng Mei , Hao He , Yahui Liu , Zhenhua Guo

Humans constantly contact objects to move and perform tasks. Thus, detecting human-object contact is important for building human-centered artificial intelligence. However, there exists no robust method to detect contact between the body…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Yixin Chen , Sai Kumar Dwivedi , Michael J. Black , Dimitrios Tzionas

Multi-object tracking (MOT) in videos remains challenging due to complex object motions and crowded scenes. Recent DETR-based frameworks offer end-to-end solutions but typically process detection and tracking queries jointly within a single…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Xu Yang , Gady Agam
‹ Prev 1 8 9 10 Next ›