English
Related papers

Related papers: ChromaGazer: Unobtrusive Visual Modulation using I…

200 papers

While users could embody virtual avatars that mirror their physical movements in Virtual Reality, these avatars' motions can be redirected to enable novel interactions. Excessive redirection, however, could break the user's sense of…

Human-Computer Interaction · Computer Science 2025-02-17 Zhipeng Li , Yishu Ji , Ruijia Chen , Tianqi Liu , Yuntao Wang , Yuanchun Shi , Yukang Yan

Medical Vision-Language Models (VLMs) often hallucinate by generating responses based on language priors rather than visual evidence, posing risks in clinical applications. We propose Visual Grounding Score Guided Decoding (VGS-Decoding), a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Govinda Kolli , Adinath Madhavrao Dukre , Behzad Bozorgtabar , Dwarikanath Mahapatra , Imran Razzak

Video chroma-lux editing, which aims to modify illumination and color while preserving structural and temporal fidelity, remains a significant challenge. Existing methods typically rely on expensive supervised training with synthetic paired…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Yifan Li , Pei Cheng , Bin Fu , Shuai Yang , Jiaying Liu

Colors are omnipresent in today's world and play a vital role in how humans perceive and interact with their surroundings. However, it is challenging for computers to imitate human color perception. This paper introduces the Human…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Pakizar Shamoi , Nuray Toganas , Muragul Muratbekova , Elnara Kadyrgali , Adilet Yerkin , Ayan Igali , Malika Ziyada , Ayana Adilova , Aron Karatayev , Yerdauit Torekhan

Recent advances in multimodal recommendation (MMR) highlight the potential of integrating visual and textual content to enrich item representations. However, existing methods often rely on coarse visual features and naive fusion strategies,…

Information Retrieval · Computer Science 2025-11-11 Hai-Dang Kieu , Min Xu , Thanh Trung Huynh , Dung D. Le

We address the problem of multi-modal object tracking in video and explore various options of fusing the complementary information conveyed by the visible (RGB) and thermal infrared (TIR) modalities including pixel-level, feature-level and…

Computer Vision and Pattern Recognition · Computer Science 2022-01-24 Zhangyong Tang , Tianyang Xu , Hui Li , Xiao-Jun Wu , Xuefeng Zhu , Josef Kittler

While current research predominantly focuses on image-based colorization, the domain of video-based colorization remains relatively unexplored. Most existing video colorization techniques operate on a frame-by-frame basis, often overlooking…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Rory Ward , Dan Bigioi , Shubhajit Basak , John G. Breslin , Peter Corcoran

Large Vision-Language Models (VLMs) often exhibit text inertia, where attention drifts from visual evidence toward linguistic priors, resulting in object hallucinations. Existing decoding strategies intervene only at the output logits and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Weijue Bu , Guan Yuan , Guixian Zhang

Vision sensors are versatile and can capture a wide range of visual cues, such as color, texture, shape, and depth. This versatility, along with the relatively inexpensive availability of machine vision cameras, played an important role in…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Muhammad Z. Alam , Zeeshan Kaleem , Sousso Kelouwani

Colour is a fundamental determinant of affective experience in immersive virtual reality (VR), yet the emotional and physiological impact of individual hues remains poorly characterised. This study investigated how fifteen calibrated…

Human-Computer Interaction · Computer Science 2025-11-19 Francesco Febbraio , Simona Collina , Christina Lepida , Panagiotis Kourtesis

Large Vision Language Models (VLMs) effectively bridge the modality gap through extensive pretraining, acquiring sophisticated visual representations aligned with language. However, it remains underexplored whether these representations,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Jiahao Guo , Sinan Du , Jingfeng Yao , Wenyu Liu , Bo Li , Haoxiang Cao , Kun Gai , Chun Yuan , Kai Wu , Xinggang Wang

Purely RGB-based vision models often fail to provide reliable cues in challenging scenarios such as nighttime and fog, leading to degraded performance and safety risks. Infrared imaging captures heat-emitting sources and provides critical…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Yuchen Guo , Junli Gong , Wenjun Dong , Yiuming Cheung , Weifeng Su

Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and construction sites significantly degrade GNSS (Global Navigation Satellite System) signal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Xiaofan Li , Zhihao Xu , Chenming Wu , Zhao Yang , Yumeng Zhang , Jiang-Jiang Liu , Haibao Yu , Fan Duan , Xiaoqing Ye , Yuan Wang , Shirui Li , Xun Sun , Ji Wan , Jun Wang

Pre-training visual and textual representations from large-scale image-text pairs is becoming a standard approach for many downstream vision-language tasks. The transformer-based models learn inter and intra-modal attention through a list…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Mohammad Abuzar Hashemi , Zhanghexuan Li , Mihir Chauhan , Yan Shen , Abhishek Satbhai , Mir Basheer Ali , Mingchen Gao , Sargur Srihari

Diffusion models have been widely used for conditional data cross-modal generation tasks such as text-to-image and text-to-video. However, state-of-the-art models still fail to align the generated visual concepts with high-level semantics…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Zizhao Hu , Shaochong Jia , Mohammad Rostami

Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust perception. However, notable limitations remain: (1) existing methods often use text only as…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Jiaqi Wu , Zhen Wang , Enhao Huang , Kangqing Shen , Yulin Wang , Yang Yue , Yifan Pu , Gao Huang

The Vision Transformer (ViT) architecture has established its place in computer vision literature, however, training ViTs for RGB-D object recognition remains an understudied topic, viewed in recent literature only through the lens of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-08 Georgios Tziafas , Hamidreza Kasaei

Vision and touch are two fundamental sensory modalities for robots, offering complementary information that enhances perception and manipulation tasks. Previous research has attempted to jointly learn visual-tactile representations to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Zhiyuan Wu , Yongqiang Zhao , Shan Luo

Image vectorization converts raster images into vector graphics composed of regions separated by curves. Typical vectorization methods first define the regions by grouping similar colored regions via color quantization, then approximate…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Roy Y. He , Sung Ha Kang , Jean-Michel Morel

Building on recent advances in video generation, generative video compression has emerged as a new paradigm for achieving visually pleasing reconstructions. However, existing methods exhibit limited exploitation of temporal correlations,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Xiaoyue Ling , Chuqin Zhou , Chunyi Li , Yunuo Chen , Yuan Tian , Guo Lu , Wenjun Zhang
‹ Prev 1 3 4 5 6 7 10 Next ›