English
Related papers

Related papers: GRAZE: Grounded Refinement and Motion-Aware Zero-S…

200 papers

Calibrating sports cameras is important for autonomous broadcasting and sports analysis. Here we propose a highly automatic method for calibrating sports cameras from a single image using synthetic data. First, we develop a novel camera…

Computer Vision and Pattern Recognition · Computer Science 2018-10-26 Jianhui Chen , James J. Little

In this paper, we introduce a Grasp Manifold Estimator (GraspME) to detect grasp affordances for objects directly in 2D camera images. To perform manipulation tasks autonomously it is crucial for robots to have such graspability models of…

Robotics · Computer Science 2021-07-06 Janik Hager , Ruben Bauer , Marc Toussaint , Jim Mainprice

We present a real-time gaze-based interaction simulation methodology using an offline dataset to evaluate the eye-tracking signal quality. This study employs three fundamental eye-movement classification algorithms to identify physiological…

Human-Computer Interaction · Computer Science 2025-05-27 Mehedi Hasan Raju , Samantha Aziz , Michael J. Proulx , Oleg V. Komogortsev

Contact-based grasp generation plays a crucial role in various applications. Recent methods typically focus on the geometric structure of objects, producing grasps with diverse hand poses and plausible contact points. However, these…

Graphics · Computer Science 2025-11-18 Zhuo Chen , Zhongqun Zhang , Yihua Cheng , Ales Leonardis , Hyung Jin Chang

Point tracking is becoming a powerful solver for motion estimation and video editing. Compared to classical feature matching, point tracking methods have the key advantage of robustly tracking points under complex camera motion trajectories…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Jianzheng Huang , Xianyu Mo , Ziling Liu , Jinyu Yang , Feng Zheng

Understanding social interactions requires reasoning over subtle non-verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with whom in multi-person videos. We introduce GRASP, a large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Junho Kim , Xu Cao , Houze Yang , Bikram Boote , Ana Jojic , Fiona Ryan , Bolin Lai , Sangmin Lee , James M. Rehg

Appearance-based gaze estimation, which uses only a regular camera to estimate human gaze, is important in various application fields. While the technique faces data bias issues, data collection protocol is often demanding, and collecting…

Human-Computer Interaction · Computer Science 2024-09-04 Mingtao Yue , Tomomi Sayuda , Miles Pennington , Yusuke Sugano

Accurate knee joint angle prediction is crucial for biomechanical analysis and rehabilitation. In this study, we introduce FocalGatedNet, a novel deep learning model that incorporates Dynamic Contextual Focus (DCF) Attention and Gated…

Robotics · Computer Science 2023-10-04 Lyes Saad Saoud , Humaid Ibrahim , Ahmad Aljarah , Irfan Hussain

Tactile sensing allows robots to gather detailed geometric information about objects through physical interaction, complementing vision-based approaches. However, efficiently acquiring useful tactile data remains challenging due to the…

Robotics · Computer Science 2026-02-27 Chung Hee Kim , Shivani Kamtikar , Tye Brady , Taskin Padir , Joshua Migdal

We study object motion path editing in videos, where the goal is to alter a target object's trajectory while preserving the original scene content. Unlike prior video editing methods that primarily manipulate appearance or rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Quynh Phung , Long Mai , Cusuh Ham , Feng Liu , Jia-Bin Huang , Aniruddha Mahapatra

We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and segmentation of any regions based on arbitrary text inputs and…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Tianhe Ren , Shilong Liu , Ailing Zeng , Jing Lin , Kunchang Li , He Cao , Jiayu Chen , Xinyu Huang , Yukang Chen , Feng Yan , Zhaoyang Zeng , Hao Zhang , Feng Li , Jie Yang , Hongyang Li , Qing Jiang , Lei Zhang

In this paper, we present Tac2Pose, an object-specific approach to tactile pose estimation from the first touch for known objects. Given the object geometry, we learn a tailored perception model in simulation that estimates a probability…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Maria Bauza , Antonia Bronars , Alberto Rodriguez

Automatically detecting graspable regions from a single depth image is a key ingredient in cloth manipulation. The large variability of cloth deformations has motivated most of the current approaches to focus on identifying specific…

Computer Vision and Pattern Recognition · Computer Science 2021-10-07 Ruijie Ren , Mohit Gurnani Rajesh , Jordi Sanchez-Riera , Fan Zhang , Yurun Tian , Antonio Agudo , Yiannis Demiris , Krystian Mikolajczyk , Francesc Moreno-Noguer

Synthetic data and novel rendering techniques have greatly influenced computer vision research in tasks like target tracking and human pose estimation. However, robotics research has lagged behind in leveraging it due to the limitations of…

Robotics · Computer Science 2024-08-23 Elia Bonetto , Chenghao Xu , Aamir Ahmad

Evaluating simultaneous localization and mapping (SLAM) algorithms necessitates high-precision and dense ground truth (GT) trajectories. But obtaining desirable GT trajectories is sometimes challenging without GT tracking sensors. As an…

Robotics · Computer Science 2023-05-23 Xiangcheng Hu , Jin Wu , Jianhao Jiao , Ruoyu Geng , Ming Liu

Grasp synthesis for 3D deformable objects remains a little-explored topic, most works aiming to minimize deformations. However, deformations are not necessarily harmful -- humans are, for example, able to exploit deformations to generate…

Robotics · Computer Science 2023-09-27 Tran Nguyen Le , Jens Lundell , Fares J. Abu-Dakka , Ville Kyrki

Grasping in cluttered scenes has always been a great challenge for robots, due to the requirement of the ability to well understand the scene and object information. Previous works usually assume that the geometry information of the objects…

Robotics · Computer Science 2021-09-28 Yiming Li , Tao Kong , Ruihang Chu , Yifeng Li , Peng Wang , Lei Li

In this paper, we study the problem of weakly-supervised temporal grounding of sentence in video. Specifically, given an untrimmed video and a query sentence, our goal is to localize a temporal segment in the video that semantically…

Computer Vision and Pattern Recognition · Computer Science 2020-01-28 Zhenfang Chen , Lin Ma , Wenhan Luo , Peng Tang , Kwan-Yee K. Wong

Visual affordance grounding aims to segment all possible interaction regions between people and objects from an image/video, which is beneficial for many applications, such as robot grasping and action recognition. However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2021-08-13 Hongchen Luo , Wei Zhai , Jing Zhang , Yang Cao , Dacheng Tao

Recently, a number of grasp detection methods have been proposed that can be used to localize robotic grasp configurations directly from sensor data without estimating object pose. The underlying idea is to treat grasp perception…

Robotics · Computer Science 2017-07-03 Andreas ten Pas , Marcus Gualtieri , Kate Saenko , Robert Platt