English
Related papers

Related papers: GRAZE: Grounded Refinement and Motion-Aware Zero-S…

200 papers

Grasping unseen objects in unconstrained, cluttered environments is an essential skill for autonomous robotic manipulation. Despite recent progress in full 6-DoF grasp learning, existing approaches often consist of complex sequential…

Robotics · Computer Science 2021-03-29 Martin Sundermeyer , Arsalan Mousavian , Rudolph Triebel , Dieter Fox

Foundation models (FM) are reshaping computer vision by reducing reliance on task-specific supervised learning and leveraging general visual representations learned at scale. In precision livestock farming, most pipelines remain dominated…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Ye Bi , Bimala Acharya , David Rosero , Juan Steibel

Early identification of hazardous actions in contact sports enables timely intervention and improves player safety. We present a method for detecting risky tackles in American football practice videos and introduce a substantially expanded…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Syed Ahsan Masud Zaidi , William Hsu , Scott Dietrich

Human motion generation is a challenging task that aims to create realistic motion imitating natural human behaviour. We focus on the well-studied behaviour of priming an object/location for pick up or put down - that is, the spotting of an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Masashi Hatano , Saptarshi Sinha , Jacob Chalk , Wei-Hong Li , Hideo Saito , Dima Damen

The pose graph is a core component of Structure-from-Motion (SfM), where images act as nodes and edges encode relative poses. Since geometric verification is expensive, SfM pipelines restrict the pose graph to a sparse set of candidate…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Tong Wei , Giorgos Tolias , Jiri Matas , Daniel Barath

Estimating the 6D pose of objects unseen during training is highly desirable yet challenging. Zero-shot object 6D pose estimation methods address this challenge by leveraging additional task-specific supervision provided by large-scale,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Andrea Caraffa , Davide Boscaini , Amir Hamza , Fabio Poiesi

Robots in the real world frequently come across identical objects in dense clutter. When evaluating grasp poses in these scenarios, a target-driven grasping system requires knowledge of spatial relations between scene objects (e.g.,…

Robotics · Computer Science 2022-03-03 Xibai Lou , Yang Yang , Changhyun Choi

Estimating the geometry level of human-scene contact aims to ground specific contact surface points at 3D human geometries, which provides a spatial prior and bridges the interaction between human and scene, supporting applications such as…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Chengfeng Wang , Wei Zhai , Yuhang Yang , Yang Cao , Zhengjun Zha

6-DoF object-agnostic grasping in unstructured environments is a critical yet challenging task in robotics. Most current works use non-optimized approaches to sample grasp locations and learn spatial features without concerning the grasping…

Robotics · Computer Science 2023-12-07 Haowen Wang , Wanhao Niu , Chungang Zhuang

Deep learning-based multi-view facial capture methods have shown impressive accuracy while being several orders of magnitude faster than a traditional mesh registration pipeline. However, the existing systems (e.g. TEMPEH) are strictly…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Jing Li , Di Kang , Zhenyu He

We propose a novel fine-grained cross-view localization method that estimates the 3 Degrees of Freedom pose of a ground-level image in an aerial image of the surroundings by matching fine-grained features between the two images. The pose is…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Zimin Xia , Alexandre Alahi

Object grasping in cluttered scenes is a widely investigated field of robot manipulation. Most of the current works focus on estimating grasp pose from point clouds based on an efficient single-shot grasp detection network. However, due to…

Robotics · Computer Science 2021-05-19 Wei Wei , Yongkang Luo , Fuyu Li , Guangyun Xu , Jun Zhong , Wanyi Li , Peng Wang

Common users have changed from mere consumers to active producers of multimedia data content. Video editing plays an important role in this scenario, calling for simple segmentation tools that can handle fast-moving and deformable video…

Computer Vision and Pattern Recognition · Computer Science 2016-06-13 Thiago Vallin Spina , Alexandre Xavier Falcão

Reconstructing physically plausible 3D human-scene interactions (HSI) from a single image currently presents a trade-off: optimization based methods offer accurate contact but are slow (~20s), while feed-forward approaches are fast yet lack…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Pradyumna YM , Yuxuan Xue , Yue Chen , Nikita Kister , István Sárándi , Gerard Pons-Moll

Proximity sensing detects an object's presence without contact. However, research has rarely explored proximity sensing in granular materials (GM) due to GM's lack of visual and complex properties. In this paper, we propose a…

Robotics · Computer Science 2023-07-19 Zeqing Zhang , Ruixing Jia , Youcan Yan , Ruihua Han , Shijie Lin , Qian Jiang , Liangjun Zhang , Jia Pan

Grasp detection methods typically target the detection of a set of free-floating hand poses that can grasp the object. However, not all of the detected grasp poses are executable due to physical constraints. Even though it is…

Robotics · Computer Science 2025-08-06 Tianyi Ko , Takuya Ikeda , Balazs Opra , Koichi Nishiwaki

Robots operating in human-centric environments require the integration of visual grounding and grasping capabilities to effectively manipulate objects based on user instructions. This work focuses on the task of referring grasp synthesis,…

Robotics · Computer Science 2023-11-13 Georgios Tziafas , Yucheng Xu , Arushi Goel , Mohammadreza Kasaei , Zhibin Li , Hamidreza Kasaei

Video analysis in tackle-collision based sports is highly subjective and exposed to bias, which is inherent in human observation, especially under time constraints. This limitation of match analysis in tackle-collision based sports can be…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Zubair Martin , Amir Patel , Sharief Hendricks

Recent video+language datasets cover domains where the interaction is highly structured, such as instructional videos, or where the interaction is scripted, such as TV shows. Both of these properties can lead to spurious cues to be…

Computer Vision and Pattern Recognition · Computer Science 2022-11-10 Alessandro Suglia , José Lopes , Emanuele Bastianelli , Andrea Vanzo , Shubham Agarwal , Malvina Nikandrou , Lu Yu , Ioannis Konstas , Verena Rieser

The choice of a grasp plays a critical role in the success of downstream manipulation tasks. Consider a task of placing an object in a cluttered scene; the majority of possible grasps may not be suitable for the desired placement. In this…

Robotics · Computer Science 2023-04-11 Zhanpeng He , Nikhil Chavan-Dafle , Jinwook Huh , Shuran Song , Volkan Isler
‹ Prev 1 2 3 10 Next ›