中文
相关论文

相关论文: SyncVP: Joint Diffusion for Synchronous Multi-Moda…

200 篇论文

Geospatial imaging leverages data from diverse sensing modalities-such as EO, SAR, and LiDAR, ranging from ground-level drones to satellite views. These heterogeneous inputs offer significant opportunities for scene understanding but…

计算机视觉与模式识别 · 计算机科学 2025-01-20 Alex Berian , Daniel Brignac , JhihYang Wu , Natnael Daba , Abhijit Mahalanobis

Spatial and temporal stream model has gained great success in video action recognition. Most existing works pay more attention to designing effective features fusion methods, which train the two-stream model in a separate way. However, it's…

计算机视觉与模式识别 · 计算机科学 2019-08-28 Jingran Zhang , Fumin Shen , Xing Xu , Heng Tao Shen

Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily focus on RGB video generation and lack the ability to support…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Guile Wu , David Huang , Dongfeng Bai , Bingbing Liu

Social media platforms serve as central hubs for content dissemination, opinion expression, and public engagement across diverse modalities. Accurately predicting the popularity of social media videos enables valuable applications in…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Liliang Ye , Yunyao Zhang , Yafeng Wu , Yi-Ping Phoebe Chen , Junqing Yu , Wei Yang , Zikai Song

Diffusion models have made significant strides in image generation, mastering tasks such as unconditional image synthesis, text-image translation, and image-to-image conversions. However, their capability falls short in the realm of video…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Gaurav Shrivastava , Abhinav Shrivastava

We propose a novel framework for the task of object-centric video prediction, i.e., extracting the compositional structure of a video sequence, as well as modeling objects dynamics and interactions from visual observations in order to…

计算机视觉与模式识别 · 计算机科学 2023-08-01 Angel Villar-Corrales , Ismail Wahdan , Sven Behnke

Purely RGB-based vision models often fail to provide reliable cues in challenging scenarios such as nighttime and fog, leading to degraded performance and safety risks. Infrared imaging captures heat-emitting sources and provides critical…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Yuchen Guo , Junli Gong , Wenjun Dong , Yiuming Cheung , Weifeng Su

This paper introduces a Multi-modal Diffusion model for Motion Prediction (MDMP) that integrates and synchronizes skeletal data and textual descriptions of actions to generate refined long-term motion predictions with quantifiable…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Leo Bringer , Joey Wilson , Kira Barton , Maani Ghaffari

The beauty of synchronized dancing lies in the synchronization of body movements among multiple dancers. While dancers utilize camera recordings for their practice, standard video interfaces do not efficiently support their activities of…

人机交互 · 计算机科学 2022-09-15 Zhongyi Zhou , Anran Xu , Koji Yatani

RGB-IR(RGB-Infrared) image pairs are frequently applied simultaneously in various applications like intelligent surveillance. However, as the number of modalities increases, the required data storage and transmission costs also double.…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Haofeng Wang , Fangtao Zhou , Qi Zhang , Zeyuan Chen , Enci Zhang , Zhao Wang , Xiaofeng Huang , Siwei Ma

Anticipating human actions is an important task that needs to be addressed for the development of reliable intelligent agents, such as self-driving cars or robot assistants. While the ability to make future predictions with high accuracy is…

计算机视觉与模式识别 · 计算机科学 2021-07-21 Olga Zatsarynna , Yazan Abu Farha , Juergen Gall

A central challenge of video prediction lies where the system has to reason the objects' future motions from image frames while simultaneously maintaining the consistency of their appearances across frames. This work introduces an…

计算机视觉与模式识别 · 计算机科学 2023-09-01 Yiqi Zhong , Luming Liang , Ilya Zharkov , Ulrich Neumann

Recent DETR-based methods have advanced the development of Video Instance Segmentation (VIS) through transformers' efficiency and capability in modeling spatial and temporal information. Despite harvesting remarkable progress, existing…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Rongkun Zheng , Lu Qi , Xi Chen , Yi Wang , Kun Wang , Yu Qiao , Hengshuang Zhao

Monocular depth estimation is a fundamental yet challenging task in computer vision, especially under complex conditions such as textureless surfaces, transparency, and specular reflections. Recent diffusion-based approaches have…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Rongjia Yu , Tong Jia , Hao Wang , Xiaofang Li , Xiao Yang , Zinuo Zhang , Cuiwei Liu

The practicality of a video surveillance system is adversely limited by the amount of queries that can be placed on human resources and their vigilance in response. To transcend this limitation, a major effort under way is to include…

计算机视觉与模式识别 · 计算机科学 2014-05-16 Samaneh Khoshrou , Jaime S. Cardoso , Luis F. Teixeira

Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine-grained motion controllability. A central limitation lies in the scarcity of such examples in…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Wonjoon Jin , Jiyun Won , Janghyeok Han , Qi Dai , Chong Luo , Seung-Hwan Baek , Sunghyun Cho

We propose a Leaked Motion Video Predictor (LMVP) to predict future frames by capturing the spatial and temporal dependencies from given inputs. The motion is modeled by a newly proposed component, motion guider, which plays the role of…

计算机视觉与模式识别 · 计算机科学 2019-06-25 Dong Wang , Yitong Li , Wei Cao , Liqun Chen , Qi Wei , Lawrence Carin

Video prediction is a useful function for autonomous driving, enabling intelligent vehicles to reliably anticipate how driving scenes will evolve and thereby supporting reasoning and safer planning. However, existing models are constrained…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Ke Li , Tianjia Yang , Kaidi Liang , Xianbiao Hu , Ruwen Qin

We present an approach for pixel-level future prediction given an input image of a scene. We observe that a scene is comprised of distinct entities that undergo motion and present an approach that operationalizes this insight. We implicitly…

计算机视觉与模式识别 · 计算机科学 2019-08-23 Yufei Ye , Maneesh Singh , Abhinav Gupta , Shubham Tulsiani

Different conditional video prediction tasks, like video future frame prediction and video frame interpolation, are normally solved by task-related models even though they share many common underlying characteristics. Furthermore, almost…

计算机视觉与模式识别 · 计算机科学 2023-04-10 Xi Ye , Guillaume-Alexandre Bilodeau