中文
相关论文

相关论文: OmniEvent: Unified Event Representation Learning

200 篇论文

Tokenizer, serving as a translator to map the intricate visual data into a compact latent space, lies at the core of visual generative models. Based on the finding that existing tokenizers are tailored to image or video inputs, this paper…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Junke Wang , Yi Jiang , Zehuan Yuan , Binyue Peng , Zuxuan Wu , Yu-Gang Jiang

Event understanding aims at understanding the content and relationship of events within texts, which covers multiple complicated information extraction tasks: event detection, event argument extraction, and event relation extraction. To…

计算与语言 · 计算机科学 2023-09-26 Hao Peng , Xiaozhi Wang , Feng Yao , Zimu Wang , Chuzhao Zhu , Kaisheng Zeng , Lei Hou , Juanzi Li

Videos convey richer information than images or text, capturing both spatial and temporal dynamics. However, most existing video customization methods rely on reference images or task-specific temporal priors, failing to fully exploit the…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Pengze Zhang , Yanze Wu , Mengtian Li , Xu Bai , Songtao Zhao , Fulong Ye , Chong Mou , Xinghui Li , Zhuowei Chen , Qian He , Mingyuan Gao

Event cameras are activity-driven bio-inspired vision sensors, thereby resulting in advantages such as sparsity,high temporal resolution, low latency, and power consumption. Given the different sensing modality of event camera and high…

计算机视觉与模式识别 · 计算机科学 2021-05-11 Lakshmi Annamalai , Vignesh Ramanathan , Chetan Singh Thakur

Event cameras have shown promise in vision applications like optical flow estimation and stereo matching, with many specialized architectures leveraging the asynchronous and sparse nature of event data. However, existing works only focus…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Pengjie Zhang , Lin Zhu , Xiao Wang , Lizhi Wang , Wanxuan Lu , Hua Huang

Representation learning, a task of learning latent vectors to represent entities, is a key task in improving search and recommender systems in web applications. Various representation learning methods have been developed, including…

This paper presents Omni-View, which extends the unified multimodal understanding and generation to 3D scenes based on multiview images, exploring the principle that "generation facilitates understanding". Consisting of understanding model,…

计算机视觉与模式识别 · 计算机科学 2026-02-02 JiaKui Hu , Shanshan Zhao , Qing-Guo Chen , Xuerui Qiu , Jialun Liu , Zhao Xu , Weihua Luo , Kaifu Zhang , Yanye Lu

The neuromorphic event cameras, which capture the optical changes of a scene, have drawn increasing attention due to their high speed and low power consumption. However, the event data are noisy, sparse, and nonuniform in the…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Chang Liu , Xiaojuan Qi , Edmund Lam , Ngai Wong

The diversity and complementarity of sensors available for Earth Observations (EO) calls for developing bespoke self-supervised multimodal learning approaches. However, current multimodal EO datasets and models typically focus on a single…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Guillaume Astruc , Nicolas Gonthier , Clement Mallet , Loic Landrieu

Event-based keypoint detection and matching holds significant potential, enabling the integration of event sensors into highly optimized Visual SLAM systems developed for frame cameras over decades of research. Unfortunately, existing…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yannick Burkhardt , Simon Schaefer , Stefan Leutenegger

This study introduces a novel approach to enhance the spatial-temporal resolution of time-event pixels based on luminance changes captured by event cameras. These cameras present unique challenges due to their low resolution and the sparse,…

图像与视频处理 · 电气工程与系统科学 2024-08-14 Waseem Shariff , Joe Lemley , Peter Corcoran

Event cameras are neuromorphic sensors that capture asynchronous and sparse event stream when per-pixel brightness changes. The state-of-the-art processing methods for event signals typically aggregate events into a frame or a grid.…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Beibei Yang , Weiling Li , Yan Fang

The field of 4D world modeling - aiming to jointly capture spatial geometry and temporal dynamics - has witnessed remarkable progress in recent years, driven by advances in large-scale generative models and multimodal learning. However, the…

Multimodal multitask learning has attracted an increasing interest in recent years. Singlemodal models have been advancing rapidly and have achieved astonishing results on various tasks across multiple domains. Multimodal learning offers…

计算机视觉与模式识别 · 计算机科学 2023-07-04 Ye Xue , Diego Klabjan , Jean Utke

In this paper, we propose EventBind, a novel and effective framework that unleashes the potential of vision-language models (VLMs) for event-based recognition to compensate for the lack of large-scale event-based datasets. In particular,…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Jiazhou Zhou , Xu Zheng , Yuanhuiyi Lyu , Lin Wang

The integration of image and event streams offers a promising approach for achieving robust visual object tracking in complex environments. However, current fusion methods achieve high performance at the cost of significant computational…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Jingjun Yang , Liangwei Fan , Jinpu Zhang , Xiangkai Lian , Hui Shen , Dewen Hu

Event cameras have gained popularity in computer vision due to their data sparsity, high dynamic range, and low latency. As a bio-inspired sensor, event cameras generate sparse and asynchronous data, which is inherently incompatible with…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Hongwei Ren , Yue Zhou , Haotian Fu , Yulong Huang , Renjing Xu , Bojun Cheng

Prior approaches injecting camera control into diffusion models have focused on specific subsets of 4D consistency tasks: novel view synthesis, text-to-video with camera control, image-to-video, amongst others. Therefore, these fragmented…

计算机视觉与模式识别 · 计算机科学 2026-01-26 Xiang Fan , Sharath Girish , Vivek Ramanujan , Chaoyang Wang , Ashkan Mirzaei , Petr Sushko , Aliaksandr Siarohin , Sergey Tulyakov , Ranjay Krishna

Modern visual agents require representations that are general, causal, and physically structured to operate in real-time streaming environments. However, current vision foundation models remain fragmented, specializing narrowly in image…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Yibin Yan , Jilan Xu , Shangzhe Di , Haoning Wu , Weidi Xie

We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model. Most previous 3D segmentation approaches are…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Wei Xu , Chunsheng Shi , Sifan Tu , Xin Zhou , Dingkang Liang , Xiang Bai
‹ 上一页 1 2 3 10 下一页 ›