English
Related papers

Related papers: THFM: A Unified Video Foundation Model for 4D Huma…

200 papers

Recently, feature upsampling has gained increasing attention owing to its effectiveness in enhancing vision foundation models (VFMs) for pixel-level understanding tasks. Existing methods typically rely on high-resolution features from the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xiaoqiong Liu , Heng Fan

Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimates dense motion but breaks down on articulated bodies, where pixel-level supervision is…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Zhanbo Huang , Xiaoming Liu , Yu Kong

We present a target-aware video diffusion model that generates videos from an input image, in which an actor interacts with a specified target while performing a desired action. The target is defined by a segmentation mask, and the action…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Taeksoo Kim , Hanbyul Joo

Video inpainting is the task of filling a region in a video in a visually convincing manner. It is very challenging due to the high dimensionality of the data and the temporal consistency required for obtaining convincing results. Recently,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Nicolas Cherel , Andrés Almansa , Yann Gousseau , Alasdair Newson

Recent advances in diffusion-based text-to-video (T2V) models have demonstrated remarkable progress, but these models still face challenges in generating videos with multiple objects. Most models struggle with accurately capturing complex…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Aimon Rahman , Jiang Liu , Ze Wang , Ximeng Sun , Jialian Wu , Xiaodong Yu , Yusheng Su , Vishal M. Patel , Zicheng Liu , Emad Barsoum

Focus is a cornerstone of photography, yet autofocus systems often fail to capture the intended subject, and users frequently wish to adjust focus after capture. We introduce a novel method for realistic post-capture refocusing using video…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 SaiKiran Tedla , Zhoutong Zhang , Xuaner Zhang , Shumian Xin

The diffusion model is widely leveraged for either video generation or video editing. As each field has its task-specific problems, it is difficult to merely develop a single diffusion for completing both tasks simultaneously. Video…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Haoyu Zhao , Tianyi Lu , Jiaxi Gu , Xing Zhang , Qingping Zheng , Zuxuan Wu , Hang Xu , Yu-Gang Jiang

Creating novel images by fusing visual cues from multiple sources is a fundamental yet underexplored problem in image-to-image generation, with broad applications in artistic creation, virtual reality and visual media. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Zeren Xiong , Yue Yu , Zedong Zhang , Shuo Chen , Jian Yang , Jun Li

While text-to-image models have achieved impressive capabilities in image generation and editing, their application across various modalities often necessitates training separate models. Inspired by existing method of single image editing…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Gihyun Kwon , Jangho Park , Jong Chul Ye

We introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue. Our core design is a progressive training approach that…

Although recent text-to-video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they still struggle to generalize to unconventional camera motions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Qiucheng Wu , Handong Zhao , Zhixin Shu , Jing Shi , Yang Zhang , Shiyu Chang

Estimating the pose of objects from images is a crucial task of 3D scene understanding, and recent approaches have shown promising results on very large benchmarks. However, these methods experience a significant performance drop when…

Computer Vision and Pattern Recognition · Computer Science 2024-10-21 Tianfu Wang , Guosheng Hu , Hongguang Wang

Human perception and understanding is a major domain of computer vision which, like many other vision subdomains recently, stands to gain from the use of large models pre-trained on large datasets. We hypothesize that the most common…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Matthieu Armando , Salma Galaaoui , Fabien Baradel , Thomas Lucas , Vincent Leroy , Romain Brégier , Philippe Weinzaepfel , Grégory Rogez

Motion prediction has been studied in different contexts with models trained on narrow distributions and applied to downstream tasks in human motion prediction and robotics. Simultaneously, recent efforts in scaling video prediction have…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Johnathan Xie , Stefan Stojanov , Cristobal Eyzaguirre , Daniel L. K. Yamins , Jiajun Wu

Vision Foundation Models (VFMs) have become the cornerstone of modern computer vision, offering robust representations across a wide array of tasks. While recent advances allow these models to handle varying input sizes during training,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Bocheng Zou , Mu Cai , Mark Stanley , Dingfu Lu , Yong Jae Lee

Generative modeling of human motion has broad applications in computer animation, virtual reality, and robotics. Conventional approaches develop separate models for different motion synthesis tasks, and typically use a model of a small size…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Jianxin Ma , Shuai Bai , Chang Zhou

In this work, we contribute to video saliency research in two ways. First, we introduce a new benchmark for predicting human eye movements during dynamic scene free-viewing, which is long-time urged in this field. Our dataset, named DHF1K…

Computer Vision and Pattern Recognition · Computer Science 2018-05-29 Wenguan Wang , Jianbing Shen , Fang Guo , Ming-Ming Cheng , Ali Borji

We introduce Home-made Diffusion Model (HDM), an efficient yet powerful text-to-image diffusion model optimized for training (and inferring) on consumer-grade hardware. HDM achieves competitive 1024x1024 generation quality while maintaining…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Shih-Ying Yeh

Synthesizing realistic animations of humans, animals, and even imaginary creatures, has long been a goal for artists and computer graphics professionals. Compared to the imaging domain, which is rich with large available datasets, the…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Sigal Raab , Inbal Leibovitch , Guy Tevet , Moab Arar , Amit H. Bermano , Daniel Cohen-Or

Multi-step prediction models, such as diffusion and rectified flow models, have emerged as state-of-the-art solutions for generation tasks. However, these models exhibit higher latency in sampling new frames compared to single-step methods.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Gaurav Shrivastava , Abhinav Shrivastava