中文
相关论文

相关论文: TD-TOG Dataset: Benchmarking Zero-Shot and One-Sho…

200 篇论文

Deep learning has achieved great success in recognizing video actions, but the collection and annotation of training data are still quite laborious, which mainly lies in two aspects: (1) the amount of required annotated data is large; (2)…

计算机视觉与模式识别 · 计算机科学 2021-11-02 Yixiong Zou , Shanghang Zhang , Guangyao Chen , Yonghong Tian , Kurt Keutzer , José M. F. Moura

Recently, general salient object detection (SOD) has made great progress with the rapid development of deep neural networks. However, task-aware SOD has hardly been studied due to the lack of task-specific datasets. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2021-05-19 Jinming Su , Changqun Xia , Jia Li

We introduce T-LESS, a new public dataset for estimating the 6D pose, i.e. translation and rotation, of texture-less rigid objects. The dataset features thirty industry-relevant objects with no significant texture and no discriminative…

计算机视觉与模式识别 · 计算机科学 2017-01-20 Tomas Hodan , Pavel Haluza , Stepan Obdrzalek , Jiri Matas , Manolis Lourakis , Xenophon Zabulis

The ability to grasp objects is an essential skill that enables many robotic manipulation tasks. Recent works have studied point cloud-based methods for object grasping by starting from simulated datasets and have shown promising…

机器人学 · 计算机科学 2022-06-07 Antonio Alliegro , Martin Rudorfer , Fabio Frattin , Aleš Leonardis , Tatiana Tommasi

Video Temporal Grounding (VTG), which aims to ground target clips from videos (such as consecutive intervals or disjoint shots) according to custom language queries (e.g., sentences or words), is key for video browsing on social media. Most…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Kevin Qinghong Lin , Pengchuan Zhang , Joya Chen , Shraman Pramanick , Difei Gao , Alex Jinpeng Wang , Rui Yan , Mike Zheng Shou

Preys in the wild evolve to be camouflaged to avoid being recognized by predators. In this way, camouflage acts as a key defence mechanism across species that is critical to survival. To detect and segment the whole scope of a camouflaged…

计算机视觉与模式识别 · 计算机科学 2023-01-04 Yunqiu Lv , Jing Zhang , Yuchao Dai , Aixuan Li , Nick Barnes , Deng-Ping Fan

Grasping is a fundamental task in robot-assisted surgery (RAS), and automating it can reduce surgeon workload while enhancing efficiency, safety, and consistency beyond teleoperated systems. Most prior approaches rely on explicit object…

机器人学 · 计算机科学 2025-08-18 Hongbin Lin , Bin Li , Kwok Wai Samuel Au

Fine-grained activity recognition enables explainable analysis of procedures for skill assessment, autonomy, and error detection in robot-assisted surgery. However, existing recognition models suffer from the limited availability of…

机器人学 · 计算机科学 2023-08-09 Kay Hutchinson , Ian Reyes , Zongyu Li , Homa Alemzadeh

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing…

机器人学 · 计算机科学 2026-02-05 Rui Tang , Guankun Wang , Long Bai , Huxin Gao , Jiewen Lai , Chi Kit Ng , Jiazheng Wang , Fan Zhang , Hongliang Ren

The ability to robustly grasp a variety of objects is essential for dexterous robots. In this paper, we present a framework for zero-shot dynamic dexterous grasping using single-view visual inputs, designed to be resilient to various…

机器人学 · 计算机科学 2025-08-15 Hui Zhang , Zijian Wu , Linyi Huang , Sammy Christen , Jie Song

Dexterous functional tool-use grasping is essential for effective robotic manipulation of tools. However, existing approaches face significant challenges in efficiently constructing large-scale datasets and ensuring generalizability to…

机器人学 · 计算机科学 2025-11-14 Sizhe Wang , Yifan Yang , Yongkang Luo , Daheng Li , Wei Wei , Yan Zhang , Peiying Hu , Yunjin Fu , Haonan Duan , Jia Sun , Peng Wang

Reliably planning fingertip grasps for multi-fingered hands lies as a key challenge for many tasks including tool use, insertion, and dexterous in-hand manipulation. This task becomes even more difficult when the robot lacks an accurate…

机器人学 · 计算机科学 2022-12-19 Martin Matak , Tucker Hermans

Temporal sentence grounding (TSG) is an important yet challenging task in multimedia information retrieval. Although previous TSG methods have achieved decent performance, they tend to capture the selection biases of frequently appeared…

计算机视觉与模式识别 · 计算机科学 2022-07-28 Daizong Liu , Xiaoye Qu , Wei Hu

Recent studies have explored pretrained (foundation) models for vision-based robotic navigation, aiming to achieve generalizable navigation and positive transfer across diverse environments while enhancing zero-shot performance in unseen…

Over the past few years, we have witnessed the success of deep learning in image recognition thanks to the availability of large-scale human-annotated datasets such as PASCAL VOC, ImageNet, and COCO. Although these datasets have covered a…

计算机视觉与模式识别 · 计算机科学 2020-12-29 Xiang Li , Tianhan Wei , Yau Pun Chen , Yu-Wing Tai , Chi-Keung Tang

Multi-object tracking and segmentation (MOTS) is a critical task for autonomous driving applications. The existing MOTS studies face two critical challenges: 1) the published datasets inadequately capture the real-world complexity for…

计算机视觉与模式识别 · 计算机科学 2021-10-18 Yiming Cui , Zhiwen Cao , Yixin Xie , Xingyu Jiang , Feng Tao , Yingjie Chen , Lin Li , Dongfang Liu

Detecting objects with visual sensors is crucial for numerous mobile robotics applications, from autonomous navigation to inspection. However, robots often need to operate under significant domains shifts from those they were trained in,…

机器人学 · 计算机科学 2026-03-20 Francesco Pasti , Riccardo De Monte , Davide Dalle Pezze , Gian Antonio Susto , Nicola Bellotto

With the human pursuit of knowledge, open-set object detection (OSOD) has been designed to identify unknown objects in a dynamic world. However, an issue with the current setting is that all the predicted unknown objects share the same…

计算机视觉与模式识别 · 计算机科学 2022-04-13 Jiyang Zheng , Weihao Li , Jie Hong , Lars Petersson , Nick Barnes

Zero-Shot Object Navigation (ZSON) requires agents to autonomously locate and approach unseen objects in unfamiliar environments and has emerged as a particularly challenging task within the domain of Embodied AI. Existing datasets for…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Ji Ma , Hongming Dai , Yao Mu , Pengying Wu , Hao Wang , Xiaowei Chi , Yang Fei , Shanghang Zhang , Chang Liu

Learning robotic grasps from visual observations is a promising yet challenging task. Recent research shows its great potential by preparing and learning from large-scale synthetic datasets. For the popular, 6 degree-of-freedom (6-DOF)…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Chaozheng Wu , Jian Chen , Qiaoyu Cao , Jianchi Zhang , Yunxin Tai , Lin Sun , Kui Jia