English
Related papers

Related papers: SAGE: Structure-Aware Generative Video Transitions…

200 papers

Human animation aims to generate temporally coherent and visually consistent videos over long sequences, yet modeling long-range dependencies while preserving frame quality remains challenging. Inspired by the human ability to leverage past…

We introduce bounded generation as a generalized task to control video generation to synthesize arbitrary camera and subject motion based only on a given start and end frame. Our objective is to fully leverage the inherent generalization…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Haiwen Feng , Zheng Ding , Zhihao Xia , Simon Niklaus , Victoria Abrevaya , Michael J. Black , Xuaner Zhang

Spatio-temporal scene graphs provide a principled representation for modeling evolving object interactions, yet existing methods remain fundamentally frame-centric: they reason only about currently visible objects, discard entities upon…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Rohith Peddi , Saurabh , Shravan Shanmugam , Likhitha Pallapothula , Yu Xiang , Parag Singla , Vibhav Gogate

Multi-modality image fusion, particularly infrared and visible, plays a crucial role in integrating diverse modalities to enhance scene understanding. Although early research prioritized visual quality, preserving fine details and adapting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Guanyao Wu , Haoyu Liu , Hongming Fu , Yichuan Peng , Jinyuan Liu , Xin Fan , Risheng Liu

State-of-the-art video generative models typically learn the distribution of video latents in the VAE space and map them to pixels using a VAE decoder. While this approach can generate high-quality videos, it suffers from slow convergence…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Jianhong Bai , Xiaoshi Wu , Xintao Wang , Xiao Fu , Yuanxing Zhang , Qinghe Wang , Xiaoyu Shi , Menghan Xia , Zuozhu Liu , Haoji Hu , Pengfei Wan , Kun Gai

Diffusion models have achieved remarkable progress in image and video stylization. However, most existing methods focus on single-style transfer, while video stylization involving multiple styles necessitates seamless transitions between…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Haoyu Zheng , Qifan Yu , Binghe Yu , Yang Dai , Wenqiao Zhang , Juncheng Li , Siliang Tang , Yueting Zhuang

Recent advancements in diffusion-based models have demonstrated significant success in generating images from text. However, video editing models have not yet reached the same level of visual quality and user control. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Ozgur Kara , Bariscan Kurtkaya , Hidir Yesiltepe , James M. Rehg , Pinar Yanardag

Neural radiance fields (NeRF) based methods have shown amazing performance in synthesizing 3D-consistent photographic images, but fail to generate multi-view portrait drawings. The key is that the basic assumption of these methods -- a…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Biao Ma , Fei Gao , Chang Jiang , Nannan Wang , Gang Xu

Video diffusion models have rich world priors, but their use in spatial tasks is limited by poor control, spatial-temporal inconsistent results, and entangled scene-camera dynamics. Current approaches, such as per-task fine-tuning or…

Graphics · Computer Science 2026-03-24 Chenxi Song , Yanming Yang , Tong Zhao , Ruibo Li , Chi Zhang

Low-latency symbolic music generation is essential for real-time improvisation and human-AI co-creation. Existing transformer-based models, however, face a trade-off between inference speed and musical quality. Traditional acceleration…

Subject-driven image generation has advanced from single- to multi-subject composition, while neglecting distinction, the ability to distinguish and generate the correct subject when inputs contain multiple candidates. This limitation…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Yuran Wang , Bohan Zeng , Chengzhuo Tong , Wenxuan Liu , Yang Shi , Xiaochen Ma , Hao Liang , Yuanxing Zhang , Wentao Zhang

Instructional video generation is an emerging task that aims to synthesize coherent demonstrations of procedural activities from textual descriptions. Such capability has broad implications for content creation, education, and human-AI…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Cheeun Hong , German Barquero , Fadime Sener , Markos Georgopoulos , Edgar Schönfeld , Stefan Popov , Yuming Du , Oscar Mañas , Albert Pumarola

Visual Place Recognition (VPR) requires robust retrieval of geotagged images despite large appearance, viewpoint, and environmental variation. Prior methods focus on descriptor fine-tuning or fixed sampling strategies yet neglect the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Shunpeng Chen , Changwei Wang , Rongtao Xu , Xingtian Pei , Yukun Song , Jinzhou Lin , Wenhao Xu , Jingyi Zhang , Li Guo , Shibiao Xu

While existing approaches excel at recognising current surgical phases, they provide limited foresight and intraoperative guidance into future procedural steps. Similarly, current anticipation methods are constrained to predicting…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Maxence Boels , Yang Liu , Prokar Dasgupta , Alejandro Granados , Sebastien Ourselin

Continuous valence-arousal estimation in real-world environments is challenging due to inconsistent modality reliability and interaction-dependent variability in audio-visual signals. Existing approaches primarily focus on modeling temporal…

Multimedia · Computer Science 2026-03-13 Yubeen Lee , Sangeun Lee , Junyeop Cha , Eunil Park

Video generation aims to produce temporally coherent sequences of visual frames, representing a pivotal advancement in Artificial Intelligence Generated Content (AIGC). Compared to static image generation, video generation poses unique…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Zhiyu Yin , Kehai Chen , Xuefeng Bai , Ruili Jiang , Juntao Li , Hongdong Li , Jin Liu , Yang Xiang , Jun Yu , Min Zhang

Retrieval-augmented question answering over heterogeneous corpora requires connected evidence across text, tables, and graph nodes. While entity-level knowledge graphs support structured access, they are costly to construct and maintain,…

Information Retrieval · Computer Science 2026-02-20 Prasham Titiya , Rohit Khoja , Tomer Wolfson , Vivek Gupta , Dan Roth

Surgical data science (SDS) is a field that analyzes patient data before, during, and after surgery to improve surgical outcomes and skills. However, surgical data is scarce, heterogeneous, and complex, which limits the applicability of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Yousef Yeganeh , Rachmadio Lazuardi , Amir Shamseddin , Emine Dari , Yash Thirani , Nassir Navab , Azade Farshad

Controllable video synthesis is a central challenge in computer vision, yet current models struggle with fine grained control beyond textual prompts, particularly for cinematic attributes like camera trajectory and genre. Existing datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zahra Dehghanian , Morteza Abolghasemi , Hamid Beigy , Hamid R. Rabiee

As humans, we are natural any-horizon reasoners, i.e., we can decide whether to iteratively skim long videos or watch short ones in full when necessary for a given task. With this in mind, one would expect video reasoning models to reason…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Jitesh Jain , Jialuo Li , Zixian Ma , Jieyu Zhang , Chris Dongjoo Kim , Sangho Lee , Rohun Tripathi , Tanmay Gupta , Christopher Clark , Humphrey Shi