English
Related papers

Related papers: Semantic-E2VID: a Semantic-Enriched Paradigm for E…

200 papers

Low-light image enhancement aims to restore the under-exposure image captured in dark scenarios. Under such scenarios, traditional frame-based cameras may fail to capture the structure and color information due to the exposure time…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Xuejian Guo , Zhiqiang Tian , Yuehang Wang , Siqi Li , Yu Jiang , Shaoyi Du , Yue Gao

Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human visual perception. While recent progress in fMRI-based image reconstruction has been notable, extending…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Minghan Yang , Lan Yang , Ke Li , Honggang Zhang , Kaiyue Pang , Yizhe Song

Despite recent advances in Video Large Language Models (Vid-LLMs), Temporal Video Grounding (TVG), which aims to precisely localize time segments corresponding to query events, remains a significant challenge. Existing methods often match…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Jiahao Nie , Wenbin An , Gongjie Zhang , Yicheng Xu , Yap-Peng Tan , Alex C. Kot , Shijian Lu

Novel view synthesis and 4D reconstruction techniques predominantly rely on RGB cameras, thereby inheriting inherent limitations such as the dependence on adequate lighting, susceptibility to motion blur, and a limited dynamic range. Event…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Chaoran Feng , Zhenyu Tang , Wangbo Yu , Yatian Pang , Yian Zhao , Jianbin Zhao , Li Yuan , Yonghong Tian

Recent years have seen a tremendous improvement in the quality of video generation and editing approaches. While several techniques focus on editing appearance, few address motion. Current approaches using text, trajectories, or bounding…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Manuel Kansy , Jacek Naruniec , Christopher Schroers , Markus Gross , Romann M. Weber

In recent years, denoising methods based on deep learning have achieved unparalleled performance at the cost of large computational complexity. In this work, we propose an Efficient Multi-stage Video Denoising algorithm, called EMVD, to…

Image and Video Processing · Electrical Eng. & Systems 2023-03-31 Matteo Maggioni , Yibin Huang , Cheng Li , Shuai Xiao , Zhongqian Fu , Fenglong Song

Segment Anything Model 2 (SAM2) shows excellent performance in video object segmentation tasks; however, the heavy computational burden hinders its application in real-time video processing. Although there have been efforts to improve the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Jing Zhang , Zhikai Li , Xuewen Liu , Qingyi Gu

With extremely high temporal resolution, event cameras have a large potential for robotics and computer vision. However, their asynchronous imaging mechanism often aggravates the measurement sensitivity to noises and brings a physical…

Computer Vision and Pattern Recognition · Computer Science 2020-07-17 Bishan Wang , Jingwei He , Lei Yu , Gui-Song Xia , Wen Yang

The Event-Enriched Image Analysis (EVENTA) Grand Challenge, hosted at ACM Multimedia 2025, introduces the first large-scale benchmark for event-level multimodal understanding. Traditional captioning and retrieval tasks largely focus on…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Thien-Phuc Tran , Minh-Quang Nguyen , Minh-Triet Tran , Tam V. Nguyen , Trong-Le Do , Duy-Nam Ly , Viet-Tham Huynh , Khanh-Duy Le , Mai-Khiem Tran , Trung-Nghia Le

Long-term temporal information is crucial for event-based perception tasks, as raw events only encode pixel brightness changes. Recent works show that when trained from scratch, recurrent models achieve better results than feedforward…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Mohammad Mohammadi , Ziyi Wu , Igor Gilitschenski

Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such as CLIP, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Yili Li , Gang Xiong , Gaopeng Gou , Xiangyan Qu , Jiamin Zhuang , Zhen Li , Junzheng Shi

In recent years, large text-to-video (T2V) synthesis models have garnered considerable attention for their abilities to generate videos from textual descriptions. However, achieving both high imaging quality and effective motion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Tongtong Su , Chengyu Wang , Bingyan Liu , Jun Huang , Dongming Lu

Event cameras are bio-inspired vision sensors that mimic retinas to asynchronously report per-pixel intensity changes rather than outputting an actual intensity image at regular intervals. This new paradigm of image sensor offers…

Computer Vision and Pattern Recognition · Computer Science 2019-04-02 Yusuke Sekikawa , Kosuke Hara , Hideo Saito

In this paper, we propose EventBind, a novel and effective framework that unleashes the potential of vision-language models (VLMs) for event-based recognition to compensate for the lack of large-scale event-based datasets. In particular,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Jiazhou Zhou , Xu Zheng , Yuanhuiyi Lyu , Lin Wang

Pixel-level Scene Understanding is one of the fundamental problems in computer vision, which aims at recognizing object classes, masks and semantics of each pixel in the given image. Since the real-world is actually video-based rather than…

Image and Video Processing · Electrical Eng. & Systems 2023-06-06 Biao Wu , Shaoli Liu , Diankai Zhang , Chengjian Zheng , Si Gao , Xiaofeng Zhang , Ning Wang

Semantic image and video segmentation stand among the most important tasks in computer vision nowadays, since they provide a complete and meaningful representation of the environment by means of a dense classification of the pixels in a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Felipe Manfio Barbosa , Fernando Santos Osório

Digital textbook (e-book) systems record student interactions with textbooks as a sequence of events called EventStream data. In the past, researchers extracted meaningful features from EventStream, and utilized them as inputs for…

Computers and Society · Computer Science 2024-07-19 Yuma Miyazaki , Valdemar Švábenský , Yuta Taniguchi , Fumiya Okubo , Tsubasa Minematsu , Atsushi Shimada

Recently, masked video modeling has been widely explored and significantly improved the model's understanding ability of visual regions at a local level. However, existing methods usually adopt random masking and follow the same…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Han Fang , Zhifei Yang , Xianghao Zang , Chao Ban , Hao Sun

Human pose estimation is critical for applications such as rehabilitation, sports analytics, and AR/VR systems. However, rapid motion and low-light conditions often introduce motion blur, significantly degrading pose estimation due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Youngho Kim , Hoonhee Cho , Kuk-Jin Yoon

Event cameras are paradigm-shifting novel sensors that report asynchronous, per-pixel brightness changes called 'events' with unparalleled low latency. This makes them ideal for high speed, high dynamic range scenes where conventional…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Timo Stoffregen , Cedric Scheerlinck , Davide Scaramuzza , Tom Drummond , Nick Barnes , Lindsay Kleeman , Robert Mahony