English
Related papers

Related papers: HuMo: Human-Centric Video Generation via Collabora…

200 papers

Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Zhe Li , Yisheng He , Lei Zhong , Weichao Shen , Qi Zuo , Lingteng Qiu , Zilong Dong , Laurence Tianruo Yang , Weihao Yuan

In natural face-to-face interaction, participants seamlessly alternate between speaking and listening, producing facial behaviors (FBs) that are finely informed by long-range context and naturally exhibit contextual appropriateness and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Xiangyu Kong , Xiaoyu Jin , Yihan Pan , Haoqin Sun , Hengde Zhu , Xiaoming Xu , Xiaoming Wei , Lu Liu , Siyang Song

Diffusion-based text-to-video generation (T2V) or image-to-video (I2V) generation have emerged as a prominent research focus. However, there exists a challenge in integrating the two generative paradigms into a unified model. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xinyu Xiao , Binbin Yang , Tingtian Li , Yipeng Yu , Sen Lei

Human motion reconstruction from monocular videos is a fundamental challenge in computer vision, with broad applications in AR/VR, robotics, and digital content creation, but remains challenging under frequent occlusions in real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Zhiyin Qian , Siwei Zhang , Bharat Lal Bhatnagar , Federica Bogo , Siyu Tang

Current deep learning results on video generation are limited while there are only a few first results on video prediction and no relevant significant results on video completion. This is due to the severe ill-posedness inherent in these…

Computer Vision and Pattern Recognition · Computer Science 2018-12-24 Haoye Cai , Chunyan Bai , Yu-Wing Tai , Chi-Keung Tang

This paper reports our solution for ACM Multimedia ViCo 2022 Conversational Head Generation Challenge, which aims to generate vivid face-to-face conversation videos based on audio and reference images. Our solution focuses on training a…

Computer Vision and Pattern Recognition · Computer Science 2022-08-03 Ailin Huang , Zhewei Huang , Shuchang Zhou

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Lijie Liu , Tianxiang Ma , Bingchuan Li , Zhuowei Chen , Jiawei Liu , Gen Li , Siyu Zhou , Qian He , Xinglong Wu

Composed Video Retrieval (CVR) is a challenging video retrieval task that utilizes multi-modal queries, consisting of a reference video and modification text, to retrieve the desired target video. The core of this task lies in understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zhiwei Chen , Yupeng Hu , Zixu Li , Zhiheng Fu , Haokun Wen , Weili Guan

Existing text-to-video (T2V) models often struggle with generating videos with sufficiently pronounced or complex actions. A key limitation lies in the text prompt's inability to precisely convey intricate motion details. To address this,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Qiang Zhou , Shaofeng Zhang , Nianzu Yang , Ye Qian , Hao Li

While modern diffusion models excel at generating high-quality and diverse images, they still struggle with high-fidelity compositional and multimodal control, particularly when users simultaneously specify text prompts, subject references,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Yusuf Dalva , Guocheng Gordon Qian , Maya Goldenberg , Tsai-Shien Chen , Kfir Aberman , Sergey Tulyakov , Pinar Yanardag , Kuan-Chieh Jackson Wang

High-quality video generation is crucial for many fields, including the film industry and autonomous driving. However, generating videos with spatiotemporal consistencies remains challenging. Current methods typically utilize attention…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Haotian Dong , Xin Wang , Di Lin , Yipeng Wu , Qin Chen , Ruonan Liu , Kairui Yang , Ping Li , Qing Guo

Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 David Romero , Ariana Bermudez , Viacheslav Iablochnikov , Hao Li , Fabio Pizzati , Ivan Laptev

Generating multi-view images from human instructions is crucial for 3D content creation. The primary challenges involve maintaining consistency across multiple views and effectively synthesizing shapes and textures under diverse conditions.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 JiaKui Hu , Yuxiao Yang , Jialun Liu , Jinbo Wu , Chen Zhao , Yanye Lu

Humanoid loco-manipulation in unstructured environments demands tight integration of egocentric perception and whole-body control. However, existing approaches either depend on external motion capture systems or fail to generalize across…

Robotics · Computer Science 2025-11-14 Shaofeng Yin , Yanjie Ze , Hong-Xing Yu , C. Karen Liu , Jiajun Wu

We address the problem of generating long-horizon videos for robotic manipulation tasks. Text-to-video diffusion models have made significant progress in photorealism, language understanding, and motion generation but struggle with…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Liudi Yang , Yang Bai , George Eskandar , Fengyi Shen , Mohammad Altillawi , Dong Chen , Soumajit Majumder , Ziyuan Liu , Gitta Kutyniok , Abhinav Valada

Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human-environment…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Haoyu Wang , Hao Tang , Donglin Di , Zhilu Zhang , Wangmeng Zuo , Feng Gao , Siwei Ma , Shiliang Zhang

Long-form video question answering requires reasoning over extended temporal contexts, making frame selection critical for large vision-language models (LVLMs) bound by finite context windows. Existing methods face a sharp trade-off:…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Dan Ben-Ami , Gabriele Serussi , Kobi Cohen , Chaim Baskin

Recent multi-modal video generation models have achieved high visual quality, but their prohibitive latency and limited temporal stability hinder real-time deployment. Streaming inference exacerbates these issues, leading to pronounced…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Rang Meng , Weipeng Wu , Yuming Li , Chenguang Ma

Creation of images using generative adversarial networks has been widely adapted into multi-modal regime with the advent of multi-modal representation models pre-trained on large corpus. Various modalities sharing a common representation…

Sound · Computer Science 2022-06-10 Yoonjeon Kim , Joel Jang , Sumin Shin

Modern video diffusion models excel at appearance synthesis but still struggle with physical consistency: objects drift, collisions lack realistic rebound, and material responses seldom match their underlying properties. We present PhyCo, a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Sriram Narayanan , Ziyu Jiang , Srinivasa Narasimhan , Manmohan Chandraker