English
Related papers

Related papers: Structure from Action: Learning Interactions for A…

200 papers

We present a dual-pathway approach for recognizing fine-grained interactions from videos. We build on the success of prior dual-stream approaches, but make a distinction between the static and dynamic representations of objects and their…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Tae Soo Kim , Jonathan Jones , Gregory D. Hager

Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are built upon vision-language models pretrained solely on 2D data,…

Robotics · Computer Science 2025-10-20 Fuhao Li , Wenxuan Song , Han Zhao , Jingbo Wang , Pengxiang Ding , Donglin Wang , Long Zeng , Haoang Li

We demonstrate the use of shape-from-shading (SfS) to improve both the quality and the robustness of 3D reconstruction of dynamic objects captured by a single camera. Unlike previous approaches that made use of SfS as a post-processing…

Computer Vision and Pattern Recognition · Computer Science 2017-08-08 Qi Liu-Yin , Rui Yu , Lourdes Agapito , Andrew Fitzgibbon , Chris Russell

Interactions play a key role in understanding objects and scenes, for both virtual and real world agents. We introduce a new general representation for proximal interactions among physical objects that is agnostic to the type of objects or…

We propose a new 3D spatial understanding task of 3D Question Answering (3D-QA). In the 3D-QA task, models receive visual information from the entire 3D scene of the rich RGB-D indoor scan and answer the given textual questions about the 3D…

Computer Vision and Pattern Recognition · Computer Science 2022-05-10 Daichi Azuma , Taiki Miyanishi , Shuhei Kurita , Motoaki Kawanabe

From refrigerators to kitchen drawers, humans interact with articulated objects effortlessly every day while completing household chores. For automating these tasks, service robots must be capable of manipulating arbitrary articulated…

Robotics · Computer Science 2026-01-06 Russell Buchanan , Adrian Röfer , João Moura , Abhinav Valada , Sethu Vijayakumar

Reconstructing the surfaces of deformable objects from correspondences between a 3D template and a 2D image is well studied under Shape-from-Template (SfT) methods; however, existing approaches break down when topological changes accompany…

Computer Vision and Pattern Recognition · Computer Science 2025-11-06 Kevin Manogue , Tomasz M Schang , Dilara Kuş , Jonas Müller , Stefan Zachow , Agniva Sengupta

Robots operating in human environments must be able to rearrange objects into semantically-meaningful configurations, even if these objects are previously unseen. In this work, we focus on the problem of building physically-valid structures…

Robotics · Computer Science 2023-04-26 Weiyu Liu , Yilun Du , Tucker Hermans , Sonia Chernova , Chris Paxton

We propose ArtiLatent, a generative framework that synthesizes human-made 3D objects with fine-grained geometry, accurate articulation, and realistic appearance. Our approach jointly models part geometry and articulation dynamics by…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Honghua Chen , Yushi Lan , Yongwei Chen , Xingang Pan

Translating high-level linguistic instructions into precise robotic actions in the physical world remains challenging, particularly when considering the feasibility of interacting with 3D objects. In this paper, we introduce 3D-TAFS, a…

Robotics · Computer Science 2025-04-08 Meng Chu , Xuan Zhang , Zhedong Zheng , Tat-Seng Chua

Perceiving and manipulating 3D articulated objects (e.g., cabinets, doors) in human environments is an important yet challenging task for future home-assistant robots. The space of 3D articulated objects is exceptionally rich in their…

Computer Vision and Pattern Recognition · Computer Science 2022-04-04 Ruihai Wu , Yan Zhao , Kaichun Mo , Zizheng Guo , Yian Wang , Tianhao Wu , Qingnan Fan , Xuelin Chen , Leonidas Guibas , Hao Dong

We present a method for inferring diverse 3D models of human-object interactions from images. Reasoning about how humans interact with objects in complex scenes from a single 2D image is a challenging task given ambiguities arising from the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Xi Wang , Gen Li , Yen-Ling Kuo , Muhammed Kocabas , Emre Aksan , Otmar Hilliges

Transparent and specular objects are frequently encountered in daily life, factories, and laboratories. However, due to the unique optical properties, the depth information on these objects is usually incomplete and inaccurate, which poses…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Yizhe Liu , Tong Jia , Da Cai , Hao Wang , Dongyue Chen

Modeling the mechanics of fluid in complex scenes is vital to applications in design, graphics, and robotics. Learning-based methods provide fast and differentiable fluid simulators, however most prior work is unable to accurately model how…

Machine Learning · Computer Science 2023-09-12 Arjun Mani , Ishaan Preetam Chandratreya , Elliot Creager , Carl Vondrick , Richard Zemel

Three-dimensional (3D) object reconstruction based on differentiable rendering (DR) is an active research topic in computer vision. DR-based methods minimize the difference between the rendered and target images by optimizing both the shape…

Computer Vision and Pattern Recognition · Computer Science 2022-11-23 Chunyu Li , Taisuke Hashimoto , Eiichi Matsumoto , Hiroharu Kato

Action Detection is a complex task that aims to detect and classify human actions in video clips. Typically, it has been addressed by processing fine-grained features extracted from a video classification backbone. Recently, thanks to the…

Computer Vision and Pattern Recognition · Computer Science 2021-03-02 Matteo Tomei , Lorenzo Baraldi , Simone Calderara , Simone Bronzin , Rita Cucchiara

General scene understanding for robotics requires flexible semantic representation, so that novel objects and structures which may not have been known at training time can be identified, segmented and grouped. We present an algorithm which…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Kirill Mazur , Edgar Sucar , Andrew J. Davison

Decomposing 3D assets into material parts is a common task for artists, yet remains a highly manual process. In this work, we introduce Select Any Material (SAMa), a material selection approach for in-the-wild objects in arbitrary 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Michael Fischer , Iliyan Georgiev , Thibault Groueix , Vladimir G. Kim , Tobias Ritschel , Valentin Deschaintre

To interact with daily-life articulated objects of diverse structures and functionalities, understanding the object parts plays a central role in both user instruction comprehension and task execution. However, the possible discordance…

Robotics · Computer Science 2024-04-02 Haoran Geng , Songlin Wei , Congyue Deng , Bokui Shen , He Wang , Leonidas Guibas

Articulated objects are pervasive in daily life. However, due to the intrinsic high-DoF structure, the joint states of the articulated objects are hard to be estimated. To model articulated objects, two kinds of shape deformations namely…

Computer Vision and Pattern Recognition · Computer Science 2021-12-15 Han Xue , Liu Liu , Wenqiang Xu , Haoyuan Fu , Cewu Lu
‹ Prev 1 4 5 6 7 8 10 Next ›