中文
相关论文

相关论文: ChromaGazer: Unobtrusive Visual Modulation using I…

200 篇论文

While users could embody virtual avatars that mirror their physical movements in Virtual Reality, these avatars' motions can be redirected to enable novel interactions. Excessive redirection, however, could break the user's sense of…

人机交互 · 计算机科学 2025-02-17 Zhipeng Li , Yishu Ji , Ruijia Chen , Tianqi Liu , Yuntao Wang , Yuanchun Shi , Yukang Yan

Medical Vision-Language Models (VLMs) often hallucinate by generating responses based on language priors rather than visual evidence, posing risks in clinical applications. We propose Visual Grounding Score Guided Decoding (VGS-Decoding), a…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Govinda Kolli , Adinath Madhavrao Dukre , Behzad Bozorgtabar , Dwarikanath Mahapatra , Imran Razzak

Video chroma-lux editing, which aims to modify illumination and color while preserving structural and temporal fidelity, remains a significant challenge. Existing methods typically rely on expensive supervised training with synthetic paired…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Yifan Li , Pei Cheng , Bin Fu , Shuai Yang , Jiaying Liu

Colors are omnipresent in today's world and play a vital role in how humans perceive and interact with their surroundings. However, it is challenging for computers to imitate human color perception. This paper introduces the Human…

Recent advances in multimodal recommendation (MMR) highlight the potential of integrating visual and textual content to enrich item representations. However, existing methods often rely on coarse visual features and naive fusion strategies,…

信息检索 · 计算机科学 2025-11-11 Hai-Dang Kieu , Min Xu , Thanh Trung Huynh , Dung D. Le

We address the problem of multi-modal object tracking in video and explore various options of fusing the complementary information conveyed by the visible (RGB) and thermal infrared (TIR) modalities including pixel-level, feature-level and…

计算机视觉与模式识别 · 计算机科学 2022-01-24 Zhangyong Tang , Tianyang Xu , Hui Li , Xiao-Jun Wu , Xuefeng Zhu , Josef Kittler

While current research predominantly focuses on image-based colorization, the domain of video-based colorization remains relatively unexplored. Most existing video colorization techniques operate on a frame-by-frame basis, often overlooking…

计算机视觉与模式识别 · 计算机科学 2024-05-10 Rory Ward , Dan Bigioi , Shubhajit Basak , John G. Breslin , Peter Corcoran

Large Vision-Language Models (VLMs) often exhibit text inertia, where attention drifts from visual evidence toward linguistic priors, resulting in object hallucinations. Existing decoding strategies intervene only at the output logits and…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Weijue Bu , Guan Yuan , Guixian Zhang

Vision sensors are versatile and can capture a wide range of visual cues, such as color, texture, shape, and depth. This versatility, along with the relatively inexpensive availability of machine vision cameras, played an important role in…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Muhammad Z. Alam , Zeeshan Kaleem , Sousso Kelouwani

Colour is a fundamental determinant of affective experience in immersive virtual reality (VR), yet the emotional and physiological impact of individual hues remains poorly characterised. This study investigated how fifteen calibrated…

人机交互 · 计算机科学 2025-11-19 Francesco Febbraio , Simona Collina , Christina Lepida , Panagiotis Kourtesis

Large Vision Language Models (VLMs) effectively bridge the modality gap through extensive pretraining, acquiring sophisticated visual representations aligned with language. However, it remains underexplored whether these representations,…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Jiahao Guo , Sinan Du , Jingfeng Yao , Wenyu Liu , Bo Li , Haoxiang Cao , Kun Gai , Chun Yuan , Kai Wu , Xinggang Wang

Purely RGB-based vision models often fail to provide reliable cues in challenging scenarios such as nighttime and fog, leading to degraded performance and safety risks. Infrared imaging captures heat-emitting sources and provides critical…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Yuchen Guo , Junli Gong , Wenjun Dong , Yiuming Cheung , Weifeng Su

Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and construction sites significantly degrade GNSS (Global Navigation Satellite System) signal…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Xiaofan Li , Zhihao Xu , Chenming Wu , Zhao Yang , Yumeng Zhang , Jiang-Jiang Liu , Haibao Yu , Fan Duan , Xiaoqing Ye , Yuan Wang , Shirui Li , Xun Sun , Ji Wan , Jun Wang

Pre-training visual and textual representations from large-scale image-text pairs is becoming a standard approach for many downstream vision-language tasks. The transformer-based models learn inter and intra-modal attention through a list…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Mohammad Abuzar Hashemi , Zhanghexuan Li , Mihir Chauhan , Yan Shen , Abhishek Satbhai , Mir Basheer Ali , Mingchen Gao , Sargur Srihari

Diffusion models have been widely used for conditional data cross-modal generation tasks such as text-to-image and text-to-video. However, state-of-the-art models still fail to align the generated visual concepts with high-level semantics…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Zizhao Hu , Shaochong Jia , Mohammad Rostami

Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust perception. However, notable limitations remain: (1) existing methods often use text only as…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Jiaqi Wu , Zhen Wang , Enhao Huang , Kangqing Shen , Yulin Wang , Yang Yue , Yifan Pu , Gao Huang

The Vision Transformer (ViT) architecture has established its place in computer vision literature, however, training ViTs for RGB-D object recognition remains an understudied topic, viewed in recent literature only through the lens of…

计算机视觉与模式识别 · 计算机科学 2023-03-08 Georgios Tziafas , Hamidreza Kasaei

Vision and touch are two fundamental sensory modalities for robots, offering complementary information that enhances perception and manipulation tasks. Previous research has attempted to jointly learn visual-tactile representations to…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Zhiyuan Wu , Yongqiang Zhao , Shan Luo

Image vectorization converts raster images into vector graphics composed of regions separated by curves. Typical vectorization methods first define the regions by grouping similar colored regions via color quantization, then approximate…

计算机视觉与模式识别 · 计算机科学 2024-09-25 Roy Y. He , Sung Ha Kang , Jean-Michel Morel

Building on recent advances in video generation, generative video compression has emerged as a new paradigm for achieving visually pleasing reconstructions. However, existing methods exhibit limited exploitation of temporal correlations,…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Xiaoyue Ling , Chuqin Zhou , Chunyi Li , Yunuo Chen , Yuan Tian , Guo Lu , Wenjun Zhang