English
Related papers

Related papers: VideoGen-of-Thought: Step-by-step generating multi…

200 papers

Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated datasets. Although…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Ruicheng Zhang , Jun Zhou , Zunnan Xu , Zihao Liu , Jiehui Huang , Mingyang Zhang , Yu Sun , Xiu Li

Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera.…

Robotics · Computer Science 2026-04-24 Songen Gu , Yuhang Zheng , Weize Li , Yupeng Zheng , Yating Feng , Xiang Li , Yilun Chen , Pengfei Li , Wenchao Ding

We propose a novel, zero-shot image generation technique called "Visual Concept Blending" that provides fine-grained control over which features from multiple reference images are transferred to a source image. If only a single reference…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Hiroya Makino , Takahiro Yamaguchi , Hiroyuki Sakai

Diffusion-based 2D virtual try-on (VTON) techniques have recently demonstrated strong performance, while the development of 3D VTON has largely lagged behind. Despite recent advances in text-guided 3D scene editing, integrating 2D VTON into…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Yukang Cao , Masoud Hadi , Liang Pan , Ziwei Liu

A method for generating narratives by analyzing single images or image sequences is presented, inspired by the time immemorial tradition of Narrative Art. The proposed method explores the multimodal capabilities of GPT-4o to interpret…

Computation and Language · Computer Science 2024-08-22 Edirlei Soares de Lima , Marco A. Casanova , Antonio L. Furtado

With the increasing popularity of autonomous driving based on the powerful and unified bird's-eye-view (BEV) representation, a demand for high-quality and large-scale multi-view video data with accurate annotation is urgently required.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Xiaofan Li , Yifu Zhang , Xiaoqing Ye

Human-Object Interaction (HOI) video reenactment with realistic motion remains a frontier in expressive digital human creation. Existing approaches primarily handle simple image-plane motion (e.g., in-plane translations), struggling with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Jinguang Tong , Jinbo Wu , Kaisiyuan Wang , Zhelun Shen , Xuan Huang , Mochu Xiang , Xuesong Li , Yingying Li , Haocheng Feng , Chen Zhao , Hang Zhou , Wei He , Chuong Nguyen , Jingdong Wang , Hongdong Li

Visual storytelling aims to generate a narrative based on a sequence of images, necessitating both vision-language alignment and coherent story generation. Most existing solutions predominantly depend on paired image-text training data,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Yuechen Wang , Wengang Zhou , Zhenbo Lu , Houqiang Li

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Junyoung Seo , Jisang Han , Jaewoo Jung , Siyoon Jin , Joungbin Lee , Takuya Narihira , Kazumi Fukuda , Takashi Shibuya , Donghoon Ahn , Shoukang Hu , Seungryong Kim , Yuki Mitsufuji

In this paper, we tackle the problem of performing consistent and unified modifications across a set of related images. This task is particularly challenging because these images may vary significantly in pose, viewpoint, and spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yue Ma , Xinyu Wang , Qianli Ma , Qinghe Wang , Mingzhe Zheng , Xiangpeng Yang , Hao Li , Chongbo Zhao , Jixuan Ying , Harry Yang , Hongyu Liu , Qifeng Chen

The advent of AI-Generated Content (AIGC) has spurred research into automated video generation to streamline conventional processes. However, automating storytelling video production, particularly for customized narratives, remains…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Panwen Hu , Jin Jiang , Jianqi Chen , Mingfei Han , Shengcai Liao , Xiaojun Chang , Xiaodan Liang

Zero-shot novel view synthesis (NVS) from a single image is an essential problem in 3D object understanding. While recent approaches that leverage pre-trained generative models can synthesize high-quality novel views from in-the-wild…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Jianglong Ye , Peng Wang , Kejie Li , Yichun Shi , Heng Wang

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Renda Li , Xiaohua Qi , Qiang Ling , Jun Yu , Ziyi Chen , Peng Chang , Mei HanJing Xiao

Subject-to-video generation has witnessed substantial progress in recent years. However, existing models still face significant challenges in faithfully following textual instructions. This limitation, commonly known as the copy-paste…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Zhuowei Chen , Bingchuan Li , Tianxiang Ma , Lijie Liu , Mingcong Liu , Yi Zhang , Gen Li , Xinghui Li , Siyu Zhou , Qian He , Xinglong Wu

Dense video captioning (DVC) aims to generate multi-sentence descriptions to elucidate the multiple events in the video, which is challenging and demands visual consistency, discoursal coherence, and linguistic diversity. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2021-11-22 Xu Yan , Zhengcong Fei , Shuhui Wang , Qingming Huang , Qi Tian

Video understanding plays a vital role in bridging low-level visual signals with high-level cognitive reasoning, and is fundamental to applications such as autonomous driving, embodied AI, and the broader pursuit of AGI. The rapid…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Yongheng Zhang , Xu Liu , Ruihan Tao , Qiguang Chen , Hao Fei , Wanxiang Che , Libo Qin

Recently, diffusion-based generative models have achieved remarkable success for image generation and edition. However, existing diffusion-based video editing approaches lack the ability to offer precise control over generated content that…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Paul Couairon , Clément Rambour , Jean-Emmanuel Haugeard , Nicolas Thome

In this work, we introduce Vision-Language Generative Pre-trained Transformer (VL-GPT), a transformer model proficient at concurrently perceiving and generating visual and linguistic data. VL-GPT achieves a unified pre-training approach for…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Jinguo Zhu , Xiaohan Ding , Yixiao Ge , Yuying Ge , Sijie Zhao , Hengshuang Zhao , Xiaohua Wang , Ying Shan

Shot transitions play a pivotal role in multi-shot video generation, as they determine the overall narrative expression and the directorial design of visual storytelling. However, recent progress has primarily focused on low-level visual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Xiaoxue Wu , Xinyuan Chen , Yaohui Wang , Yu Qiao

Despite the significant progress that has been made in video generative models, existing state-of-the-art methods can only produce videos lasting 5-16 seconds, often labeled "long-form videos". Furthermore, videos exceeding 16 seconds…

‹ Prev 1 8 9 10 Next ›