中文
相关论文

相关论文: ChromaGazer: Unobtrusive Visual Modulation using I…

200 篇论文

In the virtual realm, individuals with photosensitive epilepsy (PSE) encounter challenges when using devices, resulting in exposure to unpredictable seizure-causing visual stimuli. The current norm for preventing epileptic flashes in media…

计算机与社会 · 计算机科学 2025-07-29 Ishan Pendyala

Visual grounding aims to ground an image region through natural language, which heavily relies on cross-modal alignment. Most existing methods transfer visual/linguistic knowledge separately by fully fine-tuning uni-modal pre-trained…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Liangtao Shi , Ting Liu , Xiantao Hu , Yue Hu , Quanjun Yin , Richang Hong

An object's interior material properties, while invisible to the human eye, determine motion observed on its surface. We propose an approach that estimates heterogeneous material properties of an object from a monocular video of its surface…

计算机视觉与模式识别 · 计算机科学 2023-04-25 Berthy T. Feng , Alexander C. Ogren , Chiara Daraio , Katherine L. Bouman

Appearance-based gaze estimation has been actively studied in recent years. However, its generalization performance for unseen head poses is still a significant limitation for existing methods. This work proposes a generalizable multi-view…

计算机视觉与模式识别 · 计算机科学 2023-11-16 Yoichiro Hisadome , Tianyi Wu , Jiawei Qin , Yusuke Sugano

Interaction in virtual reality (VR) environments is essential to achieve a pleasant and immersive experience. Most of the currently existing VR applications, lack of robust object grasping and manipulation, which are the cornerstone of…

Visible and near-infrared(NIR) band sensors provide images that capture complementary spectral radiations from a scene. And the fusion of the visible and NIR image aims at utilizing their spectrum properties to enhance image quality.…

计算机视觉与模式识别 · 计算机科学 2023-09-20 Zhuo Li , Bo Li

Gaze prediction plays a critical role in Virtual Reality (VR) applications by reducing sensor-induced latency and enabling computationally demanding techniques such as foveated rendering, which rely on anticipating user attention. However,…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Christos Petrou , Harris Partaourides , Athanasios Balomenos , Yannis Kopsinis , Sotirios Chatzis

Gaze following aims to interpret human-scene interactions by predicting the person's focal point of gaze. Prevailing approaches often adopt a two-stage framework, whereby multi-modality information is extracted in the initial stage for gaze…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Yuehao Song , Xinggang Wang , Jingfeng Yao , Wenyu Liu , Jinglin Zhang , Xiangmin Xu

Gaze interaction presents a promising avenue in Virtual Reality (VR) due to its intuitive and efficient user experience. Yet, the depth control inherent in our visual system remains underutilized in current methods. In this study, we…

人机交互 · 计算机科学 2024-05-08 Chenyang Zhang , Tiansu Chen , Eric Shaffer , Elahe Soltanaghai

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as…

计算机视觉与模式识别 · 计算机科学 2021-11-01 Yi-Wen Chen , Yi-Hsuan Tsai , Ming-Hsuan Yang

Modern Visual-Aware Recommender Systems (VARS) exploit the integration of user interaction data and visual features to deliver personalized recommendations with high precision. However, their robustness against adversarial attacks remains…

信息检索 · 计算机科学 2025-07-31 Hai Ling , Tianchi Wang , Xiaohao Liu , Zhulin Tao , Lifang Yang , Xianglin Huang

Classifier-free guidance (CFG) has become a widely adopted and practical approach for enhancing generation quality and improving condition alignment. Recent studies have explored guidance mechanisms for unconditional generation, yet these…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Chaoyang Wang , Tianmeng Yang , Jingdong Wang , Yunhai Tong

Video Scene Graph Generation (VidSGG) aims to capture dynamic relationships among entities by sequentially analyzing video frames and integrating visual and semantic information. However, VidSGG is challenged by significant biases that skew…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Yanjun Li , Zhaoyang Li , Honghui Chen , Lizhi Xu

Precise color control remains a persistent failure mode in text-to-image diffusion systems, particularly in design-oriented workflows where outputs must satisfy explicit, user-specified color targets. We present an inference-time,…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Angad Singh Ahuja , Aarush Ram Anandh

It is still challenging for computer vision to imitate human color perception, e.g., color constancy, which is a fundamental perceptual ability in humans to perceive, interpret and interact with their surroundings. Among others, the…

神经元与认知 · 定量生物学 2026-04-07 Kai-Fu Yang , Dajun Xing , Yong-Jie Li

Various types of sensors can be used for Human Activity Recognition (HAR), and each of them has different strengths and weaknesses. Sometimes a single sensor cannot fully observe the user's motions from its perspective, which causes wrong…

机器学习 · 计算机科学 2024-08-05 Duc-Anh Nguyen , Cuong Pham , Nhien-An Le-Khac

Vision-Language-Action (VLA) models typically map visual observations and linguistic instructions directly to control signals. This "black-box" mapping forces a single forward pass to simultaneously handle instruction interpretation,…

机器人学 · 计算机科学 2026-05-12 Zixuan Wang , Yuxin Chen , Yuqi Liu , Jinhui Ye , Pengguang Chen , Changsheng Lu , Shu Liu , Bei Yu , Jiaya Jia

Bias in computer vision models remains a significant challenge, often resulting in unfair, unreliable, and non-generalizable AI systems. Although research into bias mitigation has intensified, progress continues to be hindered by fragmented…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Ioannis Sarridis , Christos Koutlis , Symeon Papadopoulos , Christos Diou

Camera-based adaptive traffic signal control is inherently partially observable: detections can be missed, vehicle speeds and distances can be noisy, and a phase-change decision becomes temporally irreversible once yellow onset is…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Jayawant Bodagala , Balaji Bodagala

Infrared and visible light image fusion aims to combine the strengths of both modalities to generate images that are rich in information and fulfill visual or computational requirements. This paper proposes an image fusion method based on…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Shuchen Sun , Ligen Shi , Chang Liu , Lina Wu , Jun Qiu