English
Related papers

Related papers: Observer-Actor: Active Vision Imitation Learning w…

200 papers

Free-moving object reconstruction from monocular video remains challenging, particularly without reliable pose or depth cues and under arbitrary object motion. We introduce OnlineSplatter, a novel online feed-forward framework generating…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Mark He Huang , Lin Geng Foo , Christian Theobalt , Ying Sun , De Wen Soh

While Human-Object Interaction(HOI) Detection has achieved tremendous advances in recent, it still remains challenging due to complex interactions with multiple humans and objects occurring in images, which would inevitably lead to…

Computer Vision and Pattern Recognition · Computer Science 2022-04-04 Kunlun Xu , Zhimin Li , Zhijun Zhang , Leizhen Dong , Wenhui Xu , Luxin Yan , Sheng Zhong , Xu Zou

One-shot Imitation Learning~(OSIL) aims to imbue AI agents with the ability to learn a new task from a single demonstration. To supervise the learning, OSIL typically requires a prohibitively large number of paired expert demonstrations --…

Machine Learning · Computer Science 2024-08-13 Philipp Wu , Kourosh Hakhamaneshi , Yuqing Du , Igor Mordatch , Aravind Rajeswaran , Pieter Abbeel

We present Cloth-HUGS, a Gaussian Splatting based neural rendering framework for photorealistic clothed human reconstruction that explicitly disentangles body and clothing. Unlike prior methods that absorb clothing into a single body…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Sadia Mubashshira , Nazanin Amini , Kevin Desai

Egocentric scenes exhibit frequent occlusions, varied viewpoints, and dynamic interactions compared to typical scene understanding tasks. Occlusions and varied viewpoints can lead to multi-view semantic inconsistencies, while dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Di Li , Jie Feng , Jiahao Chen , Weisheng Dong , Guanbin Li , Guangming Shi , Licheng Jiao

Attention (and distraction) recognition is a key factor in improving human-robot collaboration. We present an assembly scenario where a human operator and a cobot collaborate equally to piece together a gearbox. The setup provides multiple…

Human-Computer Interaction · Computer Science 2023-04-03 Pooja Prajod , Matteo Lavit Nicora , Matteo Malosio , Elisabeth André

In this paper, a multi-objective model-following control problem is solved using an observer-based adaptive learning scheme. The overall goal is to regulate the model-following error dynamics along with optimizing the dynamic variables of a…

Systems and Control · Electrical Eng. & Systems 2023-08-22 Mohammed I. Abouheaf , Kyriakos G. Vamvoudakis , Mohammad A. Mayyas , Hashim A. Hashim

Recent works on 3D scene understanding leverage 2D masks from visual foundation models (VFMs) to supervise radiance fields, enabling instance-level 3D segmentation. However, the supervision signals from foundation models are not…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Tsuheng Hsu , Guiyu Liu , Juho Kannala , Janne Heikkilä

Recent graph convolutional neural networks (GCNs) have shown high performance in the field of human action recognition by using human skeleton poses. However, it fails to detect human-object interaction cases successfully due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Hesham M. Shehata , Mohammad Abdolrahmani

Accurate and scalable quantification of animal pose and appearance is crucial for studying behavior. Current 3D pose estimation techniques, such as keypoint- and mesh-based techniques, often face challenges including limited…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Jack Goffinet , Youngjo Min , Carlo Tomasi , David E. Carlson

We propose ActiveSplat, an autonomous high-fidelity reconstruction system leveraging Gaussian splatting. Taking advantage of efficient and realistic rendering, the system establishes a unified framework for online mapping, viewpoint…

Robotics · Computer Science 2025-06-17 Yuetao Li , Zijia Kuang , Ting Li , Qun Hao , Zike Yan , Guyue Zhou , Shaohui Zhang

Existing methods on video-based action recognition are generally view-dependent, i.e., performing recognition from the same views seen in the training data. We present a novel multiview spatio-temporal AND-OR graph (MST-AOG) representation…

Computer Vision and Pattern Recognition · Computer Science 2014-05-14 Jiang wang , Xiaohan Nie , Yin Xia , Ying Wu , Song-Chun Zhu

Transformers have revolutionized vision and natural language processing with their ability to scale with large datasets. But in robotic manipulation, data is both limited and expensive. Can manipulation still benefit from Transformers with…

Robotics · Computer Science 2022-11-14 Mohit Shridhar , Lucas Manuelli , Dieter Fox

Accurately detecting active objects undergoing state changes is essential for comprehending human interactions and facilitating decision-making. The existing methods for active object detection (AOD) primarily rely on visual appearance of…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Dejie Yang , Yang Liu

Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent. It is difficult to…

Interactive perception enables robots to manipulate the environment and objects to bring them into states that benefit the perception process. Deformable objects pose challenges to this due to significant manipulation difficulty and…

Large-scale real-world robot data collection is a prerequisite for bringing robots into everyday deployment. However, existing pipelines often rely on specialized handheld devices to bridge the embodiment gap, which not only increases…

Robotics · Computer Science 2026-04-10 Yanwen Zou , Chenyang Shi , Wenye Yu , Han Xue , Jun Lv , Ye Pan , Chuan Wen , Cewu Lu

We introduce Gaussian Articulated Template Model GART, an explicit, efficient, and expressive representation for non-rigid articulated subject capturing and rendering from monocular videos. GART utilizes a mixture of moving 3D Gaussians to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Jiahui Lei , Yufu Wang , Georgios Pavlakos , Lingjie Liu , Kostas Daniilidis

GuessWhat?! is a two-player visual dialog guessing game where player A asks a sequence of yes/no questions (Questioner) and makes a final guess (Guesser) about a target object in an image, based on answers from player B (Oracle). Based on…

Computer Vision and Pattern Recognition · Computer Science 2021-05-26 Tao Tu , Qing Ping , Govind Thattai , Gokhan Tur , Prem Natarajan

While novel view synthesis for dynamic scenes has made significant progress, capturing skeleton models of objects and re-posing them remains a challenging task. To tackle this problem, in this paper, we propose a novel approach to…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Diwen Wan , Yuxiang Wang , Ruijie Lu , Gang Zeng