中文
相关论文

相关论文: MGTR: End-to-End Mutual Gaze Detection with Transf…

200 篇论文

Eye-tracking applications that utilize the human gaze in video understanding tasks have become increasingly important. To effectively automate the process of video analysis based on eye-tracking data, it is important to accurately replicate…

计算机视觉与模式识别 · 计算机科学 2024-04-15 Suleyman Ozdel , Yao Rong , Berat Mert Albaba , Yen-Ling Kuo , Xi Wang , Enkelejda Kasneci

Micro-gestures are unconsciously performed body gestures that can convey the emotion states of humans and start to attract more research attention in the fields of human behavior understanding and affective computing as an emerging topic.…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Zhaoqiang Xia , Hexiang Huang , Haoyu Chen , Xiaoyi Feng , Guoying Zhao

Multimodal transformer exhibits high capacity and flexibility to align image and text for visual grounding. However, the existing encoder-only grounding framework (e.g., TransVG) suffers from heavy computation due to the self-attention…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Fengyuan Shi , Ruopeng Gao , Weilin Huang , Limin Wang

Real-time multimodal inference on resource-constrained edge devices is essential for applications such as autonomous driving, human-computer interaction, and mobile health. However, prior work often overlooks the tight coupling between…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Runxi Huang , Mingxuan Yu , Mingyu Tsoi , Xiaomin Ouyang

LiDAR and camera are two important sensors for 3D object detection in autonomous driving. Despite the increasing popularity of sensor fusion in this field, the robustness against inferior image conditions, e.g., bad illumination and sensor…

计算机视觉与模式识别 · 计算机科学 2022-03-23 Xuyang Bai , Zeyu Hu , Xinge Zhu , Qingqiu Huang , Yilun Chen , Hongbo Fu , Chiew-Lan Tai

Multimodal emotion recognition identifies human emotions from various data modalities like video, text, and audio. However, we found that this task can be easily affected by noisy information that does not contain useful semantics. To this…

多媒体 · 计算机科学 2023-05-05 Yuanyuan Liu , Haoyu Zhang , Yibing Zhan , Zijing Chen , Guanghao Yin , Lin Wei , Zhe Chen

Predicting gaze behavior in virtual reality environments remains a significant challenge with implications for rendering optimization and interface design. This paper introduces a multimodal approach to VR gaze prediction that combines…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Farhaan Ebadulla , Chiraag Mudlpaur , Shreya Chaurasia , Gaurav BV

Transparent object perception is a crucial skill for applications such as robot manipulation in household and laboratory settings. Existing methods utilize RGB-D or stereo inputs to handle a subset of perception tasks including depth and…

机器人学 · 计算机科学 2023-02-24 Yi Ru Wang , Yuchi Zhao , Haoping Xu , Saggi Eppel , Alan Aspuru-Guzik , Florian Shkurti , Animesh Garg

Computer vision applications such as visual relationship detection and human object interaction can be formulated as a composite (structured) set detection problem in which both the parts (subject, object, and predicate) and the sum…

计算机视觉与模式识别 · 计算机科学 2021-08-23 Qi Dong , Zhuowen Tu , Haofu Liao , Yuting Zhang , Vijay Mahadevan , Stefano Soatto

Inspired by the fact that humans use diverse sensory organs to perceive the world, sensors with different modalities are deployed in end-to-end driving to obtain the global context of the 3D scene. In previous works, camera and LiDAR inputs…

计算机视觉与模式识别 · 计算机科学 2022-08-04 Qingwen Zhang , Mingkai Tang , Ruoyu Geng , Feiyi Chen , Ren Xin , Lujia Wang

We present a new method, called MEsh TRansfOrmer (METRO), to reconstruct 3D human pose and mesh vertices from a single image. Our method uses a transformer encoder to jointly model vertex-vertex and vertex-joint interactions, and outputs 3D…

计算机视觉与模式识别 · 计算机科学 2021-06-16 Kevin Lin , Lijuan Wang , Zicheng Liu

Real-time object detection is critical for the decision-making process for many real-world applications, such as collision avoidance and path planning in autonomous driving. This work presents an innovative real-time streaming perception…

计算机视觉与模式识别 · 计算机科学 2024-09-11 Xiang Zhang , Yufei Cui , Chenchen Fu , Weiwei Wu , Zihao Wang , Yuyang Sun , Xue Liu

Multi-view action recognition aims to recognize human actions using multiple camera views and deals with occlusion caused by obstacles or crowds. In this task, cooperation among views, which generates a joint representation by combining…

计算机视觉与模式识别 · 计算机科学 2025-11-05 Taiga Yamane , Satoshi Suzuki , Ryo Masumura , Shotaro Tora

Recognizing characters from low-resolution (LR) text images poses a significant challenge due to the information deficiency as well as the noise and blur in low-quality images. Current solutions for low-resolution text recognition (LTR)…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Hang Guo , Tao Dai , Mingyan Zhu , Guanghao Meng , Bin Chen , Zhi Wang , Shu-Tao Xia

The recently proposed end-to-end transformer detectors, such as DETR and Deformable DETR, have a cascade structure of stacking 6 decoder layers to update object queries iteratively, without which their performance degrades seriously. In…

计算机视觉与模式识别 · 计算机科学 2021-04-06 Zhuyu Yao , Jiangbo Ai , Boxun Li , Chi Zhang

Long-form video understanding, characterized by long-range temporal dependencies and multiple events, remains a challenge. Existing methods often rely on static reasoning or external visual-language models (VLMs), which face issues like…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Yuan Xie , Tianshui Chen , Zheng Ge , Lionel Ni

In this paper, we propose a transformer based approach for visual grounding. Unlike previous proposal-and-rank frameworks that rely heavily on pretrained object detectors or proposal-free frameworks that upgrade an off-the-shelf one-stage…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Ye Du , Zehua Fu , Qingjie Liu , Yunhong Wang

We address the problem of gaze target estimation, which aims to predict where a person is looking in a scene. Predicting a person's gaze target requires reasoning both about the person's appearance and the contents of the scene. Prior works…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Fiona Ryan , Ajay Bati , Sangmin Lee , Daniel Bolya , Judy Hoffman , James M. Rehg

DETR accomplishes end-to-end object detection through iteratively generating multiple object candidates based on image features and promoting one candidate for each ground-truth object. The traditional training procedure using one-to-one…

计算机视觉与模式识别 · 计算机科学 2024-01-09 Chuyang Zhao , Yifan Sun , Wenhao Wang , Qiang Chen , Errui Ding , Yi Yang , Jingdong Wang

Although current face manipulation techniques achieve impressive performance regarding quality and controllability, they are struggling to generate temporal coherent face videos. In this work, we explore to take full advantage of the…

计算机视觉与模式识别 · 计算机科学 2021-08-17 Yinglin Zheng , Jianmin Bao , Dong Chen , Ming Zeng , Fang Wen