English
Related papers

Related papers: Revelio: A Real-World Screen-Camera Communication …

200 papers

Virtual and augmented reality (VR/AR) systems are emerging technologies requiring data rates of multiple Gbps. Existing high quality VR headsets require connections through HDMI cables to a computer rendering rich graphic contents to meet…

Networking and Internet Architecture · Computer Science 2019-04-09 Mahmudur Khan , Jacob Chakareski

Spelling correction from visual input poses unique challenges for vision language models (VLMs), as it requires not only detecting but also correcting textual errors directly within images. We present ReViCo (Real Visual Correction), the…

Computation and Language · Computer Science 2025-09-23 Junhong Liang , Bojun Zhang

The event camera, benefiting from its high dynamic range and low latency, provides performance gain for low-light image enhancement. Unlike frame-based cameras, it records intensity changes with extremely high temporal resolution, capturing…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Chunyan She , Fujun Han , Chengyu Fang , Shukai Duan , Lidan Wang

Indoor service robots need perception that is robust, more privacy-friendly than RGB video, and feasible on embedded hardware. We present a camera-free 2D LiDAR object detection pipeline that encodes short-term temporal context by stacking…

Signal Processing · Electrical Eng. & Systems 2026-02-03 Soheil Behnam Roudsari , Alexandre S. Brandão , Felipe N. Martins

Capturing digital screens with smartphones frequently induces severe banding due to hardware synchronization mismatches. Existing video restoration methods struggle with these structured, periodic luminance fluctuations, often resulting in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Zhiyi Zhou , Libo Zhu , Zihan Zhou , Yulun Zhang , Xiaokang Yang

Volumetric video relighting is essential for bringing captured performances into virtual worlds, but current approaches struggle to deliver temporally stable, production-ready results. Diffusion-based intrinsic decomposition methods show…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Elisabeth Jüttner , Janelle Pfeifer , Leona Krath , Stefan Korfhage , Hannah Dröge , Matthias B. Hullin , Markus Plack

Vision-language models, such as CLIP, have achieved significant success in aligning visual and textual representations, becoming essential components of many multi-modal large language models (MLLMs) like LLaVA and OpenFlamingo. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Shizhan Gong , Yankai Jiang , Qi Dou , Farzan Farnia

Mobile robots operating indoors must be prepared to navigate challenging scenes that contain transparent surfaces. This paper proposes a novel method for the fusion of acoustic and visual sensing modalities through implicit neural…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Advaith V. Sethuraman , Onur Bagoren , Harikrishnan Seetharaman , Dalton Richardson , Joseph Taylor , Katherine A. Skinner

Video portraits relighting is critical in user-facing human photography, especially for immersive VR/AR experience. Recent advances still fail to recover consistent relit result under dynamic illuminations from monocular RGB stream,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Longwen Zhang , Qixuan Zhang , Minye Wu , Jingyi Yu , Lan Xu

Recent text-to-video diffusion transformers generate visually compelling frames, yet still struggle with temporal coherence, often producing flickering, drifting, or unstable motion. We show that these failures leave a clear imprint inside…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Nurislam Tursynbek , Zhiqiang Lao , Heather Yu , Gedas Bertasius , Marc Niethammer

Existing methods achieve high-quality facial albedo capture under controllable lighting, which increases capture cost and limits usability. We propose WildCap, a novel method for high-quality facial albedo capture from a smartphone video…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Yuxuan Han , Xin Ming , Tianxiao Li , Zhuofan Shen , Qixuan Zhang , Lan Xu , Feng Xu

While simulation tools for visible light communication (VLC) with photo detectors (PDs) have been widely investigated, similar tools for optical camera communication (OCC) with complementary metal oxide semiconductor (CMOS) sensors are…

Signal Processing · Electrical Eng. & Systems 2025-04-15 Srivathsan Chakaravarthi Narasimman , Arokiaswami Alphones

True video understanding requires making sense of non-lambertian scenes where the color of light arriving at the camera sensor encodes information about not just the last object it collided with, but about multiple mediums -- colored…

Computer Vision and Pattern Recognition · Computer Science 2019-04-05 Jean-Baptiste Alayrac , João Carreira , Andrew Zisserman

Cooperative perception enabled by Vehicle-to-Everything (V2X) communication holds significant promise for enhancing the perception capabilities of autonomous vehicles, allowing them to overcome occlusions and extend their field of view.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Hao Xiang , Zhaoliang Zheng , Xin Xia , Seth Z. Zhao , Letian Gao , Zewei Zhou , Tianhui Cai , Yun Zhang , Jiaqi Ma

Data protection methods like cryptography, despite being effective, inadvertently signal the presence of secret communication, thereby drawing undue attention. Here, we introduce an optical information hiding camera integrated with an…

Optics · Physics 2024-06-13 Bijie Bai , Ryan Lee , Yuhang Li , Tianyi Gan , Yuntian Wang , Mona Jarrahi , Aydogan Ozcan

Accurate measurement of images produced by electronic displays is critical for the evaluation of both traditional and computational displays. Traditional display measurement methods based on sparse radiometric sampling and fitting a model…

Graphics · Computer Science 2025-09-23 Yancheng Cai , Robert Wanat , Rafal Mantiuk

We introduce TransformerFusion, a transformer-based 3D scene reconstruction approach. From an input monocular RGB video, the video frames are processed by a transformer network that fuses the observations into a volumetric feature grid…

Computer Vision and Pattern Recognition · Computer Science 2021-07-07 Aljaž Božič , Pablo Palafox , Justus Thies , Angela Dai , Matthias Nießner

Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Lixin Xue , Chengwei Zheng , Georgios Paschalidis , Chen Guo , Manuel Kaufmann , Juan Zarate , Dimitrios Tzionas

Passive visible light communication (VLC) modulates light propagation or reflection to transmit data without directly modulating the light source. Thus, passive VLC provides an alternative to conventional VLC, enabling communication where…

Networking and Internet Architecture · Computer Science 2024-10-22 Yanxiang Wang , Yiran Shen , Kenuo Xu , Guangrong Zhao , Mahbub Hassan , Chenren Xu , Wen Hu

To bridge the gap between supervised semantic segmentation and real-world applications that acquires one model to recognize arbitrary new concepts, recent zero-shot segmentation attracts a lot of attention by exploring the relationships…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Quande Liu , Youpeng Wen , Jianhua Han , Chunjing Xu , Hang Xu , Xiaodan Liang