English
Related papers

Related papers: SyncMV4D: Synchronized Multi-view Joint Diffusion …

200 papers

Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine-grained motion controllability. A central limitation lies in the scarcity of such examples in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Wonjoon Jin , Jiyun Won , Janghyeok Han , Qi Dai , Chong Luo , Seung-Hwan Baek , Sunghyun Cho

Human-centric video generation has advanced rapidly, yet existing methods struggle to produce controllable and physically consistent Human-Object Interaction (HOI) videos. Existing works rely on dense control signals, template videos, or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Jiazhi Guan , Quanwei Yang , Luying Huang , Junhao Liang , Borong Liang , Haocheng Feng , Wei He , Kaisiyuan Wang , Hang Zhou , Jingdong Wang

Existing hand-object interactions (HOI) methods are largely limited to rigid objects, while 4D reconstruction methods of articulated objects generally require pre-scanning the object or even multi-view videos. It remains an unexplored but…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Zikai Wang , Zhilu Zhang , Yiqing Wang , Hui Li , Wangmeng Zuo

Most RGB-based hand-object reconstruction methods rely on object templates, while template-free methods typically assume full object visibility. This assumption often breaks in real-world settings, where fixed camera viewpoints and static…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Shibo Wang , Haonan He , Maria Parelli , Christoph Gebhardt , Zicong Fan , Jie Song

In this study, we tackle the complex task of generating 3D human-object interactions (HOI) from textual descriptions in a zero-shot text-to-3D manner. We identify and address two key challenges: the unsatisfactory outcomes of direct…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Sisi Dai , Wenhao Li , Haowen Sun , Haibin Huang , Chongyang Ma , Hui Huang , Kai Xu , Ruizhen Hu

Human-Object Interaction (HOI) recognition in videos requires understanding both visual patterns and geometric relationships as they evolve over time. Visual and geometric features offer complementary strengths. Visual features capture…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Tanqiu Qiao , Ruochen Li , Frederick W. B. Li , Yoshiki Kubotani , Shigeo Morishima , Hubert P. H. Shum

The synthesis of spatiotemporally coherent 4D content presents fundamental challenges in computer vision, requiring simultaneous modeling of high-fidelity spatial representations and physically plausible temporal dynamics. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Xiaoyan Liu , Kangrui Li , Yuehao Song , Jiaxin Liu

Recent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image-conditioned models. It is inconvenient for them to take advantage of various off-the-shelf 3D assets with multi-view…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Yanqin Jiang , Chaohui Yu , Chenjie Cao , Fan Wang , Weiming Hu , Jin Gao

Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising alternative, existing methods suffer from \textit{Representation Misalignment} due to…

To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generalization to novel objects/scenarios, and restricted accessibility -- we introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Ziyao Huang , Zixiang Zhou , Juan Cao , Yifeng Ma , Yi Chen , Zejing Rao , Zhiyong Xu , Hongmei Wang , Qin Lin , Yuan Zhou , Qinglin Lu , Fan Tang

Generating realistic 3D human-object interactions (HOIs) from text descriptions is a active research topic with potential applications in virtual and augmented reality, robotics, and animation. However, creating high-quality 3D HOIs remains…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Yixuan Zhang , Hui Yang , Chuanchen Luo , Junran Peng , Yuxi Wang , Zhaoxiang Zhang

Recent advances in 3D human-aware generation have made significant progress. However, existing methods still struggle with generating novel Human Object Interaction (HOI) from text, particularly for open-set objects. We identify three main…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Jinlu Zhang , Yixin Chen , Zan Wang , Jie Yang , Yizhou Wang , Siyuan Huang

Synthesizing accurate hands-object interactions (HOI) is critical for applications in Computer Vision, Augmented Reality (AR), and Mixed Reality (MR). Despite recent advances, the accuracy of reconstructed or generated HOI leaves room for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Théo Morales , Omid Taheri , Gerard Lacey

Photorealistic 3D full-body human reconstruction from a single image is a critical yet challenging task for applications in films and video games due to inherent ambiguities and severe self-occlusions. While recent approaches leverage SMPL…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Wenyue Chen , Peng Li , Wangguandong Zheng , Chengfeng Zhao , Mengfei Li , Yaolong Zhu , Zhiyang Dou , Ronggang Wang , Yuan Liu

We study the problem of precisely swapping objects in videos, with a focus on those interacted with by hands, given one user-provided reference object image. Despite the great advancements that diffusion models have made in video editing…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Zihui Xue , Mi Luo , Changan Chen , Kristen Grauman

We present HOIDiNi, a text-driven diffusion framework for synthesizing realistic and plausible human-object interaction (HOI). HOI generation is extremely challenging since it induces strict contact accuracies alongside a diverse motion…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Roey Ron , Guy Tevet , Haim Sawdayee , Amit H. Bermano

Hand motion plays a central role in human interaction, yet modeling realistic 4D hand motion (i.e., 3D hand pose sequences over time) remains challenging. Research in this area is typically divided into two tasks: (1) Estimation approaches…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Zhihao Sun , Tong Wu , Ruirui Tu , Daoguo Dong , Zuxuan Wu

The availability of large-scale multimodal datasets and advancements in diffusion models have significantly accelerated progress in 4D content generation. Most prior approaches rely on multiple image or video diffusion models, utilizing…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Hanwen Liang , Yuyang Yin , Dejia Xu , Hanxue Liang , Zhangyang Wang , Konstantinos N. Plataniotis , Yao Zhao , Yunchao Wei

Recent advances in diffusion-based video generation have opened new possibilities for controllable video editing, yet realistic video object insertion (VOI) remains challenging due to limited 4D scene understanding and inadequate handling…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Hoiyeong Jin , Hyojin Jang , Jeongho Kim , Junha Hyung , Kinam Kim , Dongjin Kim , Huijin Choi , Hyeonji Kim , Jaegul Choo

Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based learning. Recent methods can reconstruct temporally coherent human and object…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yubo Zhao , Yujin Chai , Yunao Dong , Chengfeng Zhao , Zijiao Zeng , Yuan Liu , Chi-Keung Tang