中文
相关论文

相关论文: FantasyHSI: Video-Generation-Centric 4D Human Synt…

200 篇论文

Web agents struggle to adapt to new websites due to the scarcity of environment specific tasks and demonstrations. Recent works have explored synthetic data generation to address this challenge, however, they suffer from data quality issues…

We present HOIDiNi, a text-driven diffusion framework for synthesizing realistic and plausible human-object interaction (HOI). HOI generation is extremely challenging since it induces strict contact accuracies alongside a diverse motion…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Roey Ron , Guy Tevet , Haim Sawdayee , Amit H. Bermano

In this paper, we study task-oriented human grasp synthesis, a new grasp synthesis task that demands both task and context awareness. At the core of our method is the task-aware contact maps. Unlike traditional contact maps that only reason…

计算机视觉与模式识别 · 计算机科学 2025-07-16 An-Lun Liu , Yu-Wei Chao , Yi-Ting Chen

Current generative models struggle to synthesize dynamic 4D driving scenes that simultaneously support temporal extrapolation and spatial novel view synthesis (NVS) without per-scene optimization. Bridging generation and novel view…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Hao Lu , Zhuang Ma , Guangfeng Jiang , Wenhang Ge , Bohan Li , Yuzhan Cai , Wenzhao Zheng , Yunpeng Zhang , Yingcong Chen

Synthesizing multi-character interactions is a challenging task due to the complex and varied interactions between the characters. In particular, precise spatiotemporal alignment between characters is required in generating close…

图形学 · 计算机科学 2022-08-05 Aman Goel , Qianhui Men , Edmond S. L. Ho

Most commodity software lacks accessible Application Programming Interfaces (APIs), requiring autonomous agents to interact solely through pixel-based Graphical User Interfaces (GUIs). In this API-free setting, large language model…

人工智能 · 计算机科学 2026-03-26 Chenwei Tang , Lin Long , Xinyu Liu , Jingyu Xing , Zizhou Wang , Joey Tianyi Zhou , Jiawei Du , Liangli Zhen , Jiancheng Lv

Common knowledge indicates that the process of constructing image datasets usually depends on the time-intensive and inefficient method of manual collection and annotation. Large models offer a solution via data generation. Nonetheless,…

计算机视觉与模式识别 · 计算机科学 2025-07-14 Haoran Sun , Haoyu Bian , Shaoning Zeng , Yunbo Rao , Xu Xu , Lin Mei , Jianping Gou

Visual interactivity understanding within visual scenes presents a significant challenge in computer vision. Existing methods focus on complex interactivities while leveraging a simple relationship model. These methods, however, struggle…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Trong-Thuan Nguyen , Pha Nguyen , Khoa Luu

Synthesizing realistic human-object interactions (HOI) in video is challenging due to the complex, instance-specific interaction dynamics of both humans and objects. Incorporating controllability in video generation further adds to the…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Wanyue Zhang , Lin Geng Foo , Thabo Beeler , Rishabh Dabral , Christian Theobalt

We introduce CFDagent, a zero-shot, multi-agent system that enables fully autonomous computational fluid dynamics (CFD) simulations from natural language prompts. CFDagent integrates three specialized LLM-driven agents: (i) the…

流体动力学 · 物理学 2025-11-13 Zhaoyue Xu , Long Wang , Chunyu Wang , Yixin Chen , Qingyong Luo , Hua-Dong Yao , Shizhao Wang , Guowei He

Speech synthesis is crucial for human-computer interaction, enabling natural and intuitive communication. However, existing datasets involve high construction costs due to manual annotation and suffer from limited character diversity,…

计算与语言 · 计算机科学 2025-04-22 Xiang Li , Duyi Pan , Hongru Xiao , Jiale Han , Jing Tang , Jiabao Ma , Wei Wang , Bo Cheng

Recent advancements in 2D and multimodal models have achieved remarkable success by leveraging large-scale training on extensive datasets. However, extending these achievements to enable free-form interactions and high-level semantic…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Shijie Zhou , Hui Ren , Yijia Weng , Shuwang Zhang , Zhen Wang , Dejia Xu , Zhiwen Fan , Suya You , Zhangyang Wang , Leonidas Guibas , Achuta Kadambi

Creating realistic characters that can react to the users' or another character's movement can benefit computer graphics, games and virtual reality hugely. However, synthesizing such reactive motions in human-human interactions is a…

图形学 · 计算机科学 2021-10-04 Qianhui Men , Hubert P. H. Shum , Edmond S. L. Ho , Howard Leung

Simulation has the potential to massively scale evaluation of self-driving systems enabling rapid development as well as safe deployment. To close the gap between simulation and the real world, we need to simulate realistic multi-agent…

机器人学 · 计算机科学 2021-01-19 Simon Suo , Sebastian Regalado , Sergio Casas , Raquel Urtasun

Understanding and generating multi-person interactions is a fundamental challenge with broad implications for robotics and social computing. While humans naturally coordinate in groups, modeling such interactions remains difficult due to…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Vongani H. Maluleke , Kie Horiuchi , Lea Wilken , Evonne Ng , Jitendra Malik , Angjoo Kanazawa

As general intelligent agents are poised for widespread deployment in diverse households, evaluation tailored to each unique unseen 3D environment has become a critical prerequisite. However, existing benchmarks suffer from severe data…

人工智能 · 计算机科学 2026-02-06 Xinyi He , Ying Yang , Chuanjian Fu , Sihan Guo , Songchun Zhu , Lifeng Fan , Zhenliang Zhang , Yujia Peng

Current validation methods often rely on recorded data and basic functional checks, which may not be sufficient to encompass the scenarios an autonomous vehicle might encounter. In addition, there is a growing need for complex scenarios…

机器人学 · 计算机科学 2024-02-08 Marc Kaufeld , Rainer Trauth , Johannes Betz

Human activity videos involve rich, varied interactions between people and objects. In this paper we develop methods for generating such videos -- making progress toward addressing the important, open problem of video generation in complex…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Megha Nawhal , Mengyao Zhai , Andreas Lehrmann , Leonid Sigal , Greg Mori

Video-based Human-Object Interaction (HOI) recognition explores the intricate dynamics between humans and objects, which are essential for a comprehensive understanding of human behavior and intentions. While previous work has made…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Tanqiu Qiao , Ruochen Li , Frederick W. B. Li , Hubert P. H. Shum

We introduce a method to synthesize animator guided human motion across 3D scenes. Given a set of sparse (3 or 4) joint locations (such as the location of a person's hand and two feet) and a seed motion sequence in a 3D scene, our method…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Aymen Mir , Xavier Puig , Angjoo Kanazawa , Gerard Pons-Moll