English
Related papers

Related papers: LG-Gaze: Learning Geometry-aware Continuous Prompt…

200 papers

Parse graphs boost human pose estimation (HPE) by integrating context and hierarchies, yet prior work mostly focuses on single modality modeling, ignoring the potential of multimodal fusion. Notably, language offers rich HPE priors like…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Shibang Liu , Xuemei Xie , Guangming Shi

Learning interpretable multimodal representations inherently relies on uncovering the conditional dependencies between heterogeneous features. However, sparse graph estimation techniques, such as Graphical Lasso (GLasso), to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Fei Wang , Yutong Zhang , Xiong Wang

This paper presents a method that utilizes multiple camera views for the gaze target estimation (GTE) task. The approach integrates information from different camera views to improve accuracy and expand applicability, addressing limitations…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Qiaomu Miao , Vivek Raju Golani , Jingyi Xu , Progga Paromita Dutta , Minh Hoai , Dimitris Samaras

We propose a two-stage multimodal framework that enhances disease classification and region-aware radiology report generation from chest X-rays, leveraging the MIMIC-Eye dataset. In the first stage, we introduce a gaze-guided contrastive…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Tanjim Islam Riju , Shuchismita Anwar , Saman Sarker Joy , Farig Sadeque , Swakkhar Shatabda

In recent years, the accuracy of gaze estimation techniques has gradually improved, but existing methods often rely on large datasets or large models to improve performance, which leads to high demands on computational resources. In terms…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Zhang Cheng , Yanxia Wang , Guoyu Xia

We present recurrent geometry-aware neural networks that integrate visual information across multiple views of a scene into 3D latent feature tensors, while maintaining an one-to-one mapping between 3D physical locations in the world scene…

Computer Vision and Pattern Recognition · Computer Science 2018-11-15 Ricson Cheng , Ziyan Wang , Katerina Fragkiadaki

Domain Generalized person Re-identification (DG Re-ID) is a challenging task, where models are trained on source domains but tested on unseen target domains. Although previous pure vision-based models have achieved significant progress, the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Jiachen Li , Xiaojin Gong , Dongping Zhang

The complex application scenarios have raised critical requirements for precise and generalizable gaze estimation methods. Recently, the pre-trained CLIP has achieved remarkable performance on various vision tasks, but its potentials have…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Lin Zhang , Yi Tian , XiYun Wang , Wanru Xu , Yi Jin , Yaping Huang

Recent years have witnessed remarkable progress in the development of large vision-language models (LVLMs). Benefiting from the strong language backbones and efficient cross-modal alignment strategies, LVLMs exhibit surprising capabilities…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Zejun Li , Ye Wang , Mengfei Du , Qingwen Liu , Binhao Wu , Jiwen Zhang , Chengxing Zhou , Zhihao Fan , Jie Fu , Jingjing Chen , Xuanjing Huang , Zhongyu Wei

As an indicator of human attention gaze is a subtle behavioral cue which can be exploited in many applications. However, inferring 3D gaze direction is challenging even for deep neural networks given the lack of large amount of data…

Computer Vision and Pattern Recognition · Computer Science 2019-04-25 Yu Yu , Gang Liu , Jean-Marc Odobez

Video-based gaze estimation methods aim to capture the inherently temporal dynamics of human eye gaze from multiple image frames. However, since models must capture both spatial and temporal relationships, performance is limited by the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Alexandre Personnic , Mihai Bâce

While pretrained language models (PLMs) have been shown to possess a plethora of linguistic knowledge, the existing body of research has largely neglected extralinguistic knowledge, which is generally difficult to obtain by pretraining on…

Computation and Language · Computer Science 2024-01-30 Valentin Hofmann , Goran Glavaš , Nikola Ljubešić , Janet B. Pierrehumbert , Hinrich Schütze

The academic field of learning instruction-guided visual navigation can be generally categorized into high-level category-specific search and low-level language-guided navigation, depending on the granularity of language instruction, in…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Gengze Zhou , Yicong Hong , Zun Wang , Chongyang Zhao , Mohit Bansal , Qi Wu

Human gaze is essential for various appealing applications. Aiming at more accurate gaze estimation, a series of recent works propose to utilize face and eye images simultaneously. Nevertheless, face and eye images only serve as independent…

Computer Vision and Pattern Recognition · Computer Science 2020-01-03 Yihua Cheng , Shiyao Huang , Fei Wang , Chen Qian , Feng Lu

Class-incremental learning requires a learning system to continually learn knowledge of new classes and meanwhile try to preserve previously learned knowledge of old classes. As current state-of-the-art methods based on Vision-Language…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Jiantao Tan , Peixian Ma , Tong Yu , Wentao Zhang , Ruixuan Wang

Multimodal Large Language Models (MLLMs) have demonstrated impressive progress in single-image grounding and general multi-image understanding. Recently, some methods begin to address multi-image grounding. However, they are constrained by…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Shurong Zheng , Yousong Zhu , Hongyin Zhao , Fan Yang , Yufei Zhan , Ming Tang , Jinqiao Wang

Continual learning is conventionally tackled through sequential fine-tuning, a process that, while enabling adaptation, inherently favors plasticity over the stability needed to retain prior knowledge. While existing approaches attempt to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Ghada Sokar , Gintare Karolina Dziugaite , Anurag Arnab , Ahmet Iscen , Pablo Samuel Castro , Cordelia Schmid

Gaze correction aims to redirect the person's gaze into the camera by manipulating the eye region, and it can be considered as a specific image resynthesis problem. Gaze correction has a wide range of applications in real life, such as…

Computer Vision and Pattern Recognition · Computer Science 2019-06-04 Jichao Zhang , Meng Sun , Jingjing Chen , Hao Tang , Yan Yan , Xueying Qin , Nicu Sebe

Linear attention methods offer a compelling alternative to softmax attention due to their efficiency in recurrent decoding. Recent research has focused on enhancing standard linear attention by incorporating gating while retaining its…

Machine Learning · Computer Science 2025-04-08 Yingcong Li , Davoud Ataee Tarzanagh , Ankit Singh Rawat , Maryam Fazel , Samet Oymak

Despite their outstanding performance, large language models (LLMs) suffer notorious flaws related to their preference for simple, surface-level textual relations over full semantic complexity of the problem. This proposal investigates a…

Computation and Language · Computer Science 2022-06-20 Michal Štefánik
‹ Prev 1 8 9 10 Next ›