English
Related papers

Related papers: Open-Vocabulary Functional 3D Human-Scene Interact…

200 papers

Continual learning refers to the ability of humans and animals to incrementally learn over time in a given environment. Trying to simulate this learning process in machines is a challenging task, also due to the inherent difficulty in…

Computer Vision and Pattern Recognition · Computer Science 2021-09-17 Enrico Meloni , Alessandro Betti , Lapo Faggi , Simone Marullo , Matteo Tiezzi , Stefano Melacci

Traditionally, 3D scene synthesis requires expert knowledge and significant manual effort. Automating this process could greatly benefit fields such as architectural design, robotics simulation, virtual reality, and gaming. Recent…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Rui Huang , Guangyao Zhai , Zuria Bauer , Marc Pollefeys , Federico Tombari , Leonidas Guibas , Gao Huang , Francis Engelmann

High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present a high-fidelity, real-time conversational digital human…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Hongbin Huang , Junwei Li , Tianxin Xie , Zhuang Li , Cekai Weng , Yaodong Yang , Yue Luo , Li Liu , Jing Tang , Zhijing Shao , Zeyu Wang

Humans interact with objects all the time. Enabling a humanoid to learn human-object interaction (HOI) is a key step for future smart animation and intelligent robotics systems. However, recent progress in physics-based HOI requires…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Yinhuai Wang , Jing Lin , Ailing Zeng , Zhengyi Luo , Jian Zhang , Lei Zhang

Recent works learn 3D representation explicitly under text-3D guidance. However, limited text-3D data restricts the vocabulary scale and text control of generations. Generators may easily fall into a stereotype concept for certain text…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Tianyu Huang , Yihan Zeng , Bowen Dong , Hang Xu , Songcen Xu , Rynson W. H. Lau , Wangmeng Zuo

While recent 3D generative models can produce high-quality texture images, they often fail to capture human preferences or meet task-specific requirements. Moreover, a core challenge in the 3D texture generation domain is that most existing…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 AmirHossein Zamani , Tianhao Xie , Amir G. Aghdam , Tiberiu Popa , Eugene Belilovsky

Realistic visual simulations are omnipresent, yet their creation requires computing time, rendering, and expert animation knowledge. Open-vocabulary visual effects generation from text inputs emerges as a promising solution that can unlock…

Graphics · Computer Science 2026-01-01 Luca Collorone , Mert Kiray , Indro Spinelli , Fabio Galasso , Benjamin Busam

In recent years, there has been a significant amount of research on algorithms and control methods for distributed collaborative robots. However, the emergence of collective behavior in a swarm is still difficult to predict and control.…

Robotics · Computer Science 2024-08-21 Pengming Zhu , Zhiwen Zeng , Weijia Yao , Wei Dai , Huimin Lu , Zongtan Zhou

Traditional 3D scene understanding approaches rely on labeled 3D datasets to train a model for a single task with supervision. We propose OpenScene, an alternative approach where a model predicts dense features for 3D scene points that are…

Computer Vision and Pattern Recognition · Computer Science 2023-04-07 Songyou Peng , Kyle Genova , Chiyu "Max" Jiang , Andrea Tagliasacchi , Marc Pollefeys , Thomas Funkhouser

Achieving unified 3D perception and reasoning across tasks such as segmentation, retrieval, and relation understanding remains challenging, as existing methods are either object-centric or rely on costly training for inter-object reasoning.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Yaxu Xie , Abdalla Arafa , Alireza Javanmardi , Christen Millerdurai , Jia Cheng Hu , Shaoxiang Wang , Alain Pagani , Didier Stricker

Synthesizing high-quality 3D face models from natural language descriptions is very valuable for many applications, including avatar creation, virtual reality, and telepresence. However, little research ever tapped into this task. We argue…

Computer Vision and Pattern Recognition · Computer Science 2023-05-08 Menghua Wu , Hao Zhu , Linjia Huang , Yiyu Zhuang , Yuanxun Lu , Xun Cao

Enabling mobile robots to perform long-term tasks in dynamic real-world environments is a formidable challenge, especially when the environment changes frequently due to human-robot interactions or the robot's own actions. Traditional…

Robotics · Computer Science 2025-03-20 Zhijie Yan , Shufei Li , Zuoxu Wang , Lixiu Wu , Han Wang , Jun Zhu , Lijiang Chen , Jihong Liu

Realistic 3D indoor scene synthesis is vital for embodied AI and digital content creation. It can be naturally divided into two subtasks: object generation and layout generation. While recent generative models have significantly advanced…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Xingjian Ran , Yixuan Li , Linning Xu , Mulin Yu , Bo Dai

Grasping and manipulating objects is an important human skill. Since most objects are designed to be manipulated by human hands, anthropomorphic hands can enable richer human-robot interaction. Desirable grasps are not only stable, but also…

Robotics · Computer Science 2019-07-26 Samarth Brahmbhatt , Ankur Handa , James Hays , Dieter Fox

Achieving natural dyadic interaction requires generating facial expressions that are emotionally appropriate and socially aligned with human preference. Human feedback offers a compelling mechanism to guide such alignment, yet how to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Xu Chen , Rui Gao , Xinjie Zhang , Haoyu Zhang , Che Sun , Zhi Gao , Yuwei Wu , Yunde Jia

Intelligent agents must autonomously interact with the environments to perform daily tasks based on human-level instructions. They need a foundational understanding of the world to accurately interpret these instructions, along with precise…

Artificial Intelligence · Computer Science 2025-08-22 Zhen Wu , Jiaman Li , Pei Xu , C. Karen Liu

Building 3D maps of the environment is central to robot navigation, planning, and interaction with objects in a scene. Most existing approaches that integrate semantic concepts with 3D maps largely remain confined to the closed-set setting:…

Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explicitly represented. An…

Robotics · Computer Science 2026-02-17 Huajian Zeng , Lingyun Chen , Jiaqi Yang , Yuantai Zhang , Fan Shi , Peidong Liu , Xingxing Zuo

To enable more natural face-to-face interactions, conversational agents need to adapt their behavior to their interlocutors. One key aspect of this is generation of appropriate non-verbal behavior for the agent, for example facial gestures,…

Computer Vision and Pattern Recognition · Computer Science 2020-10-26 Patrik Jonell , Taras Kucherenko , Gustav Eje Henter , Jonas Beskow

Modern 3D object detection datasets are constrained by narrow class taxonomies and costly manual annotations, limiting their ability to scale to open-world settings. In contrast, 2D vision-language models trained on web-scale image-text…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Atharv Goel , Mehar Khurana