English
Related papers

Related papers: Generative Event Pretraining with Foundation Model…

200 papers

Edge vision systems combining sensing and embedded processing promise low-latency, decentralized, and energy-efficient solutions that forgo reliance on the cloud. As opposed to conventional frame-based vision sensors, event-based cameras…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Yufeng Yang , Adrian Kneip , Charlotte Frenkel

Event cameras have emerged as a promising sensing modality for autonomous navigation systems, owing to their high temporal resolution, high dynamic range and negligible motion blur. To process the asynchronous temporal event streams from…

Machine Learning · Computer Science 2024-03-26 Shrihari Sridharan , Surya Selvam , Kaushik Roy , Anand Raghunathan

Self-supervised learning has emerged as a highly effective approach in the fields of natural language processing and computer vision. It is also applicable to brain signals such as electroencephalography (EEG) data, given the abundance of…

Signal Processing · Electrical Eng. & Systems 2024-01-22 Yuqi Chen , Kan Ren , Kaitao Song , Yansen Wang , Yifan Wang , Dongsheng Li , Lili Qiu

The ever-increasing demands for intuitive interactions in Virtual Reality has triggered a boom in the realm of Facial Expression Recognition (FER). To address the limitations in existing approaches (e.g., narrow receptive fields and…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Yande Li , Mingjie Wang , Minglun Gong , Yonggang Lu , Li Liu

In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers (ViTs) designed for multimodal large language models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Yan Fang , Mengcheng Lan , Zilong Huang , Weixian Lei , Yunqing Zhao , Yujie Zhong , Yingchen Yu , Qi She , Yao Zhao , Yunchao Wei

We present Recurrent Vision Transformers (RVTs), a novel backbone for object detection with event cameras. Event cameras provide visual information with sub-millisecond latency at a high-dynamic range and with strong robustness against…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Mathias Gehrig , Davide Scaramuzza

Foundation model-based semantic transmission has recently shown great potential in wireless image communication. However, existing methods exhibit two major limitations: (i) they overlook the varying importance of semantic components for…

Image and Video Processing · Electrical Eng. & Systems 2025-09-30 Fangyu Liu , Peiwen Jiang , Wenjin Wang , Chao-Kai Wen , Shi Jin , Jun Zhang

Reliable detection and classification of power system events are critical for maintaining grid stability and situational awareness. Existing approaches often depend on limited labeled datasets, which restricts their ability to generalize to…

Signal Processing · Electrical Eng. & Systems 2026-05-22 Yi Hu , Zheyuan Cheng

Across domains, metrics and measurements are fundamental to identifying challenges, informing decisions, and resolving conflicts. Despite the abundance of data available in this information age, not only can it be challenging for a single…

Software Engineering · Computer Science 2024-10-02 Ti-Chung Cheng , Carmen Badea , Christian Bird , Thomas Zimmermann , Robert DeLine , Nicole Forsgren , Denae Ford

Event cameras offer unique advantages for facial keypoint alignment under challenging conditions, such as low light and rapid motion, due to their high temporal resolution and robustness to varying illumination. However, existing RGB facial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Donghwa Kang , Junho Kim , Dongwoo Kang

Aligning video generative models with human preferences remains challenging: current approaches rely on Vision-Language Models (VLMs) for reward modeling, but these models struggle to capture subtle temporal dynamics. We propose a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Shivanshu Shekhar , Uttaran Bhattacharya , Raghavendra Addanki , Mehrab Tanjim , Somdeb Sarkhel , Tong Zhang

Event-based semantic segmentation (ESS) is a fundamental yet challenging task for event camera sensing. The difficulties in interpreting and annotating event data limit its scalability. While domain adaptation from images to event data can…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Lingdong Kong , Youquan Liu , Lai Xing Ng , Benoit R. Cottereau , Wei Tsang Ooi

Multi-modal Event Reasoning (MMER) endeavors to endow machines with the ability to comprehend intricate event relations across diverse data modalities. MMER is fundamental and underlies a wide broad of applications. Despite extensive…

Artificial Intelligence · Computer Science 2024-04-17 Zhengwei Tao , Zhi Jin , Junqiang Huang , Xiancai Chen , Xiaoying Bai , Haiyan Zhao , Yifan Zhang , Chongyang Tao

Vision foundation models (VFMs) are predominantly developed using data-centric methods. These methods require training on vast amounts of data usually with high-quality labels, which poses a bottleneck for most institutions that lack both…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Jiabo Huang , Chen Chen , Lingjuan Lyu

The event camera's low power consumption and ability to capture microsecond brightness changes make it attractive for various computer vision tasks. Existing event representation methods typically convert events into frames, voxel grids, or…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Bin Jiang , Zhihao Li , M. Salman Asif , Xun Cao , Zhan Ma

Most existing real-time deep models trained with each frame independently may produce inconsistent results across the temporal axis when tested on a video sequence. A few methods take the correlations in the video sequence into…

Computer Vision and Pattern Recognition · Computer Science 2022-02-28 Yifan Liu , Chunhua Shen , Changqian Yu , Jingdong Wang

Pretext training followed by task-specific fine-tuning has been a successful approach in vision and language domains. This paper proposes a self-supervised pretext training framework tailored to event sequence data. We introduce a novel…

Machine Learning · Computer Science 2024-02-19 Yimu Wang , He Zhao , Ruizhi Deng , Frederick Tung , Greg Mori

Event-based vision sensors offer high time resolution, high dynamic range, and low power consumption, yet event-based vision models lag behind conventional frame-based vision methods. We argue that this gap is partly due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Jens Egholm Pedersen , Dimitris Korakovounis , Jörg Conradt

The broad scope of obstacle avoidance has led to many kinds of computer vision-based approaches. Despite its popularity, it is not a solved problem. Traditional computer vision techniques using cameras and depth sensors often focus on…

Computer Vision and Pattern Recognition · Computer Science 2022-12-06 Celyn Walters , Simon Hadfield

The event camera, benefiting from its high dynamic range and low latency, provides performance gain for low-light image enhancement. Unlike frame-based cameras, it records intensity changes with extremely high temporal resolution, capturing…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Chunyan She , Fujun Han , Chengyu Fang , Shukai Duan , Lidan Wang