English
Related papers

Related papers: GRAZE: Grounded Refinement and Motion-Aware Zero-S…

200 papers

High fidelity digital 3D environments have been proposed in recent years, however, it remains extremely challenging to automatically equip such environment with realistic human bodies. Existing work utilizes images, depth or semantic maps…

Computer Vision and Pattern Recognition · Computer Science 2020-11-13 Siwei Zhang , Yan Zhang , Qianli Ma , Michael J. Black , Siyu Tang

Task-oriented grasping (TOG), which refers to synthesizing grasps on an object that are configurationally compatible with the downstream manipulation task, is the first milestone towards tool manipulation. Analogous to the activation of two…

Robotics · Computer Science 2024-10-10 Chao Tang , Dehao Huang , Wenlong Dong , Ruinian Xu , Hong Zhang

We introduce AO-Grasp, a grasp proposal method that generates 6 DoF grasps that enable robots to interact with articulated objects, such as opening and closing cabinets and appliances. AO-Grasp consists of two main contributions: the…

In Scene Graph Generation (SGG), structured representations are extracted from visual inputs as object nodes and connecting predicates, enabling image-based reasoning for diverse downstream tasks. While fully supervised SGG has improved…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Abdelrahman Elskhawy , Mengze Li , Nassir Navab , Benjamin Busam

The segmentation of a gaze trace into its constituent eye movements has been actively researched since the early days of eye tracking. As we move towards more naturalistic viewing conditions, the segmentation becomes even more challenging…

Multimedia · Computer Science 2019-12-11 Ioannis Agtzidis , Mikhail Startsev , Michael Dorr

Grasping for novel objects is important for robot manipulation in unstructured environments. Most of current works require a grasp sampling process to obtain grasp candidates, combined with local feature extractor using deep learning. This…

Robotics · Computer Science 2020-03-24 Peiyuan Ni , Wenguang Zhang , Xiaoxiao Zhu , Qixin Cao

Image understanding is a foundational task in computer vision, with recent applications emerging in soccer posture analysis. However, existing publicly available datasets lack comprehensive information, notably in the form of posture…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Calvin Yeung , Kenjiro Ide , Keisuke Fujii

Fatigue monitoring is central in association football due to its links with injury risk and tactical performance. However, objective fatigue-related indicators are commonly derived from subjective self-reported metrics, biomarkers derived…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Xavier Bou , Nathan Correger , Alexandre Cloots , Cédric Gavage , Silvio Giancola , Cédric Schwartz , François Delvaux , Rudi Cloots , Marc Van Droogenbroeck , Anthony Cioppa

Dynamic environments that include unstructured moving objects pose a hard problem for Simultaneous Localization and Mapping (SLAM) performance. The motion of rigid objects can be typically tracked by exploiting their texture and geometric…

Computer Vision and Pattern Recognition · Computer Science 2021-08-03 Huayan Zhang , Tianwei Zhang , Tin Lun Lam , Sethu Vijayakumar

We evaluate the video understanding capabilities of existing foundation models (FMs) using a carefully designed experiment protocol consisting of three hallmark tasks (action recognition,temporal localization, and spatiotemporal…

This paper introduces Grounding DINO 1.5, a suite of advanced open-set object detection models developed by IDEA Research, which aims to advance the "Edge" of open-set object detection. The suite encompasses two models: Grounding DINO 1.5…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Tianhe Ren , Qing Jiang , Shilong Liu , Zhaoyang Zeng , Wenlong Liu , Han Gao , Hongjie Huang , Zhengyu Ma , Xiaoke Jiang , Yihao Chen , Yuda Xiong , Hao Zhang , Feng Li , Peijun Tang , Kent Yu , Lei Zhang

This work explores the performance of a large video understanding foundation model on the downstream task of human fall detection on untrimmed video and leverages a pretrained vision transformer for multi-class action detection, with…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Till Grutschus , Ola Karrar , Emir Esenov , Ekta Vats

Grasping objects in cluttered scenarios is a challenging task in robotics. Performing pre-grasp actions such as pushing and shifting to scatter objects is a way to reduce clutter. Based on deep reinforcement learning, we propose a…

Robotics · Computer Science 2021-07-07 Dafa Ren , Xiaoqiang Ren , Xiaofan Wang , S. Tejaswi Digumarti , Guodong Shi

We're interested in the problem of estimating object states from touch during manipulation under occlusions. In this work, we address the problem of estimating object poses from touch during planar pushing. Vision-based tactile sensors…

Robotics · Computer Science 2021-03-30 Paloma Sodhi , Michael Kaess , Mustafa Mukadam , Stuart Anderson

We present the first approach to build hierarchical task-driven 3D scene graphs of arbitrary indoor or outdoor environments using an uncalibrated monocular camera in real-time. We leverage geometric foundation models to estimate geometric…

Robotics · Computer Science 2026-05-26 Dominic Maggio , Nicolas Gorlo , Luca Carlone

Despite the recent development of learning-based gaze estimation methods, most methods require one or more eye or face region crops as inputs and produce a gaze direction vector as output. Cropping results in a higher resolution in the eye…

Computer Vision and Pattern Recognition · Computer Science 2023-05-10 Haldun Balim , Seonwook Park , Xi Wang , Xucong Zhang , Otmar Hilliges

Training computers to understand, model, and synthesize human grasping requires a rich dataset containing complex 3D object shapes, detailed contact information, hand pose and shape, and the 3D body motion over time. While "grasping" is…

Computer Vision and Pattern Recognition · Computer Science 2020-08-26 Omid Taheri , Nima Ghorbani , Michael J. Black , Dimitrios Tzionas

Rationale discovery is defined as finding a subset of the input data that maximally supports the prediction of downstream tasks. In the context of graph machine learning, graph rationale is defined to locate the critical subgraph in the…

Machine Learning · Computer Science 2025-01-28 Zhe Xu , Menghai Pan , Yuzhong Chen , Huiyuan Chen , Yuchen Yan , Mahashweta Das , Hanghang Tong

The 3D Human Pose Estimation (3D HPE) task uses 2D images or videos to predict human joint coordinates in 3D space. Despite recent advancements in deep learning-based methods, they mostly ignore the capability of coupling accessible texts…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Jinglin Xu , Yijie Guo , Yuxin Peng

Modelling the trajectorial motion of humans along the ground is a foundational task in the quantitative analysis of sports like association football. Most existing models of football player motion have not been validated yet with respect to…

Other Computer Science · Computer Science 2022-05-02 M. Renkin , J. Bischofberger , E. Schikuta , A. Baca
‹ Prev 1 8 9 10 Next ›