English
Related papers

Related papers: Generative Event Pretraining with Foundation Model…

200 papers

High-speed vision sensing is essential for real-time perception in applications such as robotics, autonomous vehicles, and industrial automation. Traditional frame-based vision systems suffer from motion blur, high latency, and redundant…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Riadul Islam , Joey Mulé , Dhandeep Challagundla , Shahmir Rizvi , Sean Carson

Event temporal graphs have been shown as convenient and effective representations of complex temporal relations between events in text. Recent studies, which employ pre-trained language models to auto-regressively generate linearised graphs…

Computation and Language · Computer Science 2024-04-03 Xingwei Tan , Yuxiang Zhou , Gabriele Pergola , Yulan He

Complex Event Processing (CEP) is an event processing paradigm to perform real-time analytics over streaming data and match high-level event patterns. Presently, CEP is limited to process structured data stream. Video streams are…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Piyush Yadav , Dhaval Salwala , Edward Curry

Generative models face a fundamental challenge: they must simultaneously learn high-level semantic concepts (what to generate) and low-level synthesis details (how to generate it). Conventional end-to-end training entangles these distinct,…

Machine Learning · Computer Science 2025-09-30 Deyuan Liu , Peng Sun , Xufeng Li , Tao Lin

In autonomous driving, relying solely on frame-based cameras can lead to inaccuracies caused by factors like long exposure times, high-speed motion, and challenging lighting conditions. To address these issues, we introduce a bio-inspired…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Hu Cao , Jiong Liu , Xingzhuo Yan , Rui Song , Yan Xia , Walter Zimmer , Guang Chen , Alois Knoll

Event cameras excel at high-speed, low-power, and high-dynamic-range scene perception. However, as they fundamentally record only relative intensity changes rather than absolute intensity, the resulting data streams suffer from a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Gang Xu , Zhiyu Zhu , Junhui Hou

While language models have become impactful in many real-world applications, video generation remains largely confined to entertainment. Motivated by video's inherent capacity to demonstrate physical-world information that is difficult to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Junhao Cheng , Liang Hou , Xin Tao , Jing Liao

Event Cameras, also known as Neuromorphic sensors, capture changes in local light intensity at the pixel level, producing asynchronously generated data termed ``events''. This distinct data format mitigates common issues observed in…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Khadija Iddrisu , Waseem Shariff , Noel E. OConnor , Joseph Lemley , Suzanne Little

As neuromorphic sensors, event cameras asynchronously record changes in brightness as streams of sparse events with the advantages of high temporal resolution and high dynamic range. Reconstructing intensity images from events is a highly…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Weilun Li , Lei Sun , Ruixi Gao , Qi Jiang , Yuqin Ma , Kaiwei Wang , Ming-Hsuan Yang , Luc Van Gool , Danda Pani Paudel

Prediction over event sequences is critical for many real-world applications in Information Retrieval and Natural Language Processing. Future Event Generation (FEG) is a challenging task in event sequence prediction because it requires not…

Computation and Language · Computer Science 2022-08-19 Li Lin , Yixin Cao , Lifu Huang , Shu'ang Li , Xuming Hu , Lijie Wen , Jianmin Wang

Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a significant gap remains between the high-level semantic focus…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Ruowen Zhao , Bangguo Li , Zuyan Liu , Yinan Liang , Junliang Ye , Fangfu Liu , Diankun Wu , Zhengyi Wang , Xumin Yu , Yongming Rao , Han Hu , Jun Zhu

Forecasting from partial observations is central to world modeling. Many recent methods represent the world through images, and reduce forecasting to stochastic video generation. Although such methods excel at realism and visual fidelity,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Gabrijel Boduljak , Yushi Lan , Christian Rupprecht , Andrea Vedaldi

By pretraining to synthesize coherent images from perturbed inputs, generative models inherently learn to understand object boundaries and scene compositions. How can we repurpose these generative representations for general-purpose…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Om Khangaonkar , Hamed Pirsiavash

Event cameras offer promising advantages such as high dynamic range and low latency, making them well-suited for challenging lighting conditions and fast-moving scenarios. However, reconstructing 3D scenes from raw event streams is…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Jiaxu Wang , Junhao He , Ziyi Zhang , Mingyuan Sun , Jingkai Sun , Renjing Xu

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li

This study explores the application of self-supervised learning techniques for event sequences. It is a key modality in various applications such as banking, e-commerce, and healthcare. However, there is limited research on self-supervised…

Event-stream representation is the first step for many computer vision tasks using event cameras. It converts the asynchronous event-streams into a formatted structure so that conventional machine learning models can be applied easily.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Qiang Qu , Xiaoming Chen , Yuk Ying Chung , Yiran Shen

Audio-visual event parsing plays a crucial role in understanding multimodal video content, but existing methods typically rely on offline processing of entire videos with huge model sizes, limiting their real-time applicability. We…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Xiao Yu , Yan Fang , Xiaojie Jin , Yao Zhao , Yunchao Wei

Large Language Model (LLM)-based Vision-Language Models (VLMs) have substantially extended the boundaries of visual understanding capabilities. However, their high computational demands hinder deployment on resource-constrained edge…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Haotong Qin , Cheng Hu , Michele Magno

Event-based cameras provide accurate and high temporal resolution measurements for performing computer vision tasks in challenging scenarios, such as high-dynamic range environments and fast-motion maneuvers. Despite their advantages,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Mohammad Rostami , Dayuan Jian , Ruitong Sun