English
Related papers

Related papers: HOIGen-1M: A Large-scale Dataset for Human-Object …

200 papers

The efficacy of video generation models heavily depends on the quality of their training datasets. Most previous video generation models are trained on short video clips, while recently there has been increasing interest in training long…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Tianwei Xiong , Yuqing Wang , Daquan Zhou , Zhijie Lin , Jiashi Feng , Xihui Liu

Modeling 4D human-object interaction (HOI) is a compelling challenge in computer vision and an essential technology powering virtual and mixed-reality applications. While existing works have achieved promising results on specific HOI…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Mengfei Zhang , Jinlu Zhang , Zhigang Tu

We propose CG-HOI, the first method to address the task of generating dynamic 3D human-object interactions (HOIs) from text. We model the motion of both human and object in an interdependent fashion, as semantically rich human motion rarely…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Christian Diller , Angela Dai

Text-Video Retrieval plays an important role in multi-modal understanding and has attracted increasing attention in recent years. Most existing methods focus on constructing contrastive pairs between whole videos and complete caption…

Computer Vision and Pattern Recognition · Computer Science 2022-12-20 Jie Jiang , Shaobo Min , Weijie Kong , Dihong Gong , Hongfa Wang , Zhifeng Li , Wei Liu

Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, image, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Liyang Chen , Tianxiang Ma , Jiawei Liu , Bingchuan Li , Zhuowei Chen , Lijie Liu , Xu He , Gen Li , Qian He , Zhiyong Wu

We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from…

Computer Vision and Pattern Recognition · Computer Science 2022-09-30 Uriel Singer , Adam Polyak , Thomas Hayes , Xi Yin , Jie An , Songyang Zhang , Qiyuan Hu , Harry Yang , Oron Ashual , Oran Gafni , Devi Parikh , Sonal Gupta , Yaniv Taigman

Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this area. However, existing public datasets are primarily composed…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Yiming Ju , Jijin Hu , Zhengxiong Luo , Haoge Deng , hanyu Zhao , Li Du , Chengwei Wu , Donglin Hao , Xinlong Wang , Tengfei Pan

Text-to-image (T2I) generation models have significantly advanced in recent years. However, effective interaction with these models is challenging for average users due to the need for specialized prompt engineering knowledge and the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Minbin Huang , Yanxin Long , Xinchi Deng , Ruihang Chu , Jiangfeng Xiong , Xiaodan Liang , Hong Cheng , Qinglin Lu , Wei Liu

Recent great advances in video generation models have demonstrated their potential to produce high-quality videos, bringing challenges to effective evaluation. Unlike human evaluation, existing automated evaluation metrics lack highlevel…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Zhun Mou , Bin Xia , Zhengchao Huang , Wenming Yang , Jiaya Jia

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Youliang Zhang , Zhengguang Zhou , Zhentao Yu , Ziyao Huang , Teng Hu , Sen Liang , Guozhen Zhang , Ziqiao Peng , Shunkai Li , Yi Chen , Zixiang Zhou , Yuan Zhou , Qinglin Lu , Xiu Li

In this paper, we develop \textbf{MP-HOI}, a powerful Multi-modal Prompt-based HOI detector designed to leverage both textual descriptions for open-set generalization and visual exemplars for handling high ambiguity in descriptions,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Jie Yang , Bingliang Li , Ailing Zeng , Lei Zhang , Ruimao Zhang

Existing studies on talking video generation have predominantly focused on single-person monologues or isolated facial animations, limiting their applicability to realistic multi-human interactions. To bridge this gap, we introduce MIT, a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Zeyu Zhu , Weijia Wu , Mike Zheng Shou

With the rapid development of generative models, Artificial Intelligence-Generated Contents (AIGC) have exponentially increased in daily lives. Among them, Text-to-Video (T2V) generation has received widespread attention. Though many T2V…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Tengchuan Kou , Xiaohong Liu , Zicheng Zhang , Chunyi Li , Haoning Wu , Xiongkuo Min , Guangtao Zhai , Ning Liu

Recent advances in Large Multimodal Models (LMMs) have expanded their capabilities to video understanding, with Text-to-Video (T2V) models excelling in generating videos from textual prompts. However, they still frequently produce…

Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on a vocabulary of interactions at training and inference time,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Francesco Tonini , Alessandro Conti , Lorenzo Vaquero , Cigdem Beyan , Elisa Ricci

Recent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Haopeng Fang , Di Qiu , Binjie Mao , He Tang

Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a significant challenge. In response to this challenge, we present…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xiangqing Zheng , Chengyue Wu , Kehai Chen , Min Zhang

While large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to dataset limitations. Existing datasets often lack…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Sirui Xu , Dongting Li , Yucheng Zhang , Xiyan Xu , Qi Long , Ziyin Wang , Yunzhi Lu , Shuchang Dong , Hezi Jiang , Akshat Gupta , Yu-Xiong Wang , Liang-Yan Gui

Human-object interaction (HOI) synthesis is important for various applications, ranging from virtual reality to robotics. However, acquiring 3D HOI data is challenging due to its complexity and high cost, limiting existing methods to the…

Graphics · Computer Science 2025-03-27 Yuke Lou , Yiming Wang , Zhen Wu , Rui Zhao , Wenjia Wang , Mingyi Shi , Taku Komura

Understanding and synthesizing realistic 3D hand-object interactions (HOI) is critical for applications ranging from immersive AR/VR to dexterous robotics. Existing methods struggle with generalization, performing well on closed-set objects…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Zhenhao Zhang , Ye Shi , Lingxiao Yang , Suting Ni , Qi Ye , Jingya Wang