English
Related papers

Related papers: Zero-shot Interactive Perception

200 papers

Zero-Shot Learning (ZSL) aims to recognize unseen classes by generalizing the knowledge, i.e., visual and semantic relationships, obtained from seen classes, where image augmentation techniques are commonly applied to improve the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-02 Zhi Chen , Pengfei Zhang , Jingjing Li , Sen Wang , Zi Huang

Large language models (LLMs) have been effectively used for many computer vision tasks, including image classification. In this paper, we present a simple yet effective approach for zero-shot image classification using multimodal LLMs.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Abdelrahman Abdelhamed , Mahmoud Afifi , Alec Go

In this paper, we study the problem of Compositional Zero-Shot Learning (CZSL), which is to recognize novel attribute-object combinations with pre-existing concepts. Recent researchers focus on applying large-scale Vision-Language…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Zhaoheng Zheng , Haidong Zhu , Ram Nevatia

We present SLIP (SAM+CLIP), an enhanced architecture for zero-shot object segmentation. SLIP combines the Segment Anything Model (SAM) \cite{kirillov2023segment} with the Contrastive Language-Image Pretraining (CLIP)…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Saaketh Koundinya Gundavarapu , Arushi Arora , Shreya Agarwal

We study the task of zero-shot vision-and-language navigation (ZS-VLN), a practical yet challenging problem in which an agent learns to navigate following a path described by language instructions without requiring any path-instruction…

Computer Vision and Pattern Recognition · Computer Science 2023-08-17 Peihao Chen , Xinyu Sun , Hongyan Zhi , Runhao Zeng , Thomas H. Li , Gaowen Liu , Mingkui Tan , Chuang Gan

Remarkable progress has been made in recent years in the fields of vision, language, and robotics. We now have vision models capable of recognizing objects based on language queries, navigation systems that can effectively control mobile…

Robotics · Computer Science 2024-11-20 Peiqi Liu , Yaswanth Orru , Jay Vakil , Chris Paxton , Nur Muhammad Mahi Shafiullah , Lerrel Pinto

Industrial robot applications require increasingly flexible systems that non-expert users can easily adapt for varying tasks and environments. However, different adaptations benefit from different interaction modalities. We present an…

Current facial expression recognition (FER) models are often designed in a supervised learning manner and thus are constrained by the lack of large-scale facial expression images with high-quality annotations. Consequently, these models…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Zengqun Zhao , Yu Cao , Shaogang Gong , Ioannis Patras

Large-scale vision-language models such as CLIP have achieved remarkable success in zero-shot image recognition, yet their predictions remain largely opaque to human understanding. In contrast, Concept Bottleneck Models provide…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Onat Ozdemir , Anders Christensen , Stephan Alaniz , Zeynep Akata , Emre Akbas

Vision-language models (VLMs) like CLIP (Contrastive Language-Image Pre-Training) have seen remarkable success in visual recognition, highlighting the increasing need to safeguard the intellectual property (IP) of well-trained models.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Lianyu Wang , Meng Wang , Huazhu Fu , Daoqiang Zhang

The Zero-Shot Object Navigation (ZSON) task requires embodied agents to find a previously unseen object by navigating in unfamiliar environments. Such a goal-oriented exploration heavily relies on the ability to perceive, understand, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Linqing Zhong , Chen Gao , Zihan Ding , Yue Liao , Huimin Ma , Shifeng Zhang , Xu Zhou , Si Liu

Task-oriented grasping of unfamiliar objects is a necessary skill for robots in dynamic in-home environments. Inspired by the human capability to grasp such objects through intuition about their shape and structure, we present a novel…

Robotics · Computer Science 2024-03-28 Samuel Li , Sarthak Bhagat , Joseph Campbell , Yaqi Xie , Woojun Kim , Katia Sycara , Simon Stepputtis

Learning accurate models of the physical world is required for a lot of robotic manipulation tasks. However, during manipulation, robots are expected to interact with unknown workpieces so that building predictive models which can…

Machine Learning · Computer Science 2020-11-03 Wenyu Zhang , Skyler Seto , Devesh K. Jha

While Vision-Language Models (VLMs) are set to transform robotic navigation, existing methods often underutilize their reasoning capabilities. To unlock the full potential of VLMs in robotics, we shift their role from passive observers to…

Robotics · Computer Science 2025-11-13 Mobin Habibpour , Fatemeh Afghah

Zero-shot learning is a learning regime that recognizes unseen classes by generalizing the visual-semantic relationship learned from the seen classes. To obtain an effective ZSL model, one may resort to curating training samples from…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Zhi Chen , Yadan Luo , Sen Wang , Jingjing Li , Zi Huang

Effective Human-Robot Interaction (HRI) is crucial for future service robots in aging societies. Existing solutions are biased toward only well-trained objects, creating a gap when dealing with new objects. Currently, HRI systems using…

Robotics · Computer Science 2025-03-13 Yuzhi Lai , Shenghai Yuan , Youssef Nassar , Mingyu Fan , Thomas Weber , Matthias Rätsch

Zero-shot human skeleton-based action recognition aims to construct a model that can recognize actions outside the categories seen during training. Previous research has focused on aligning sequences' visual and semantic spatial…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Haojun Xu , Yan Gao , Jie Li , Xinbo Gao

Zero-shot learning (ZSL) aims to recognize a set of unseen classes without any training images. The standard approach to ZSL requires a set of training images annotated with seen class labels and a semantic descriptor for seen/unseen…

Computer Vision and Pattern Recognition · Computer Science 2019-03-19 Nanyi Fei , Jiechao Guan , Zhiwu Lu , Tao Xiang , Ji-Rong Wen

Intelligent virtual assistants are currently designed to perform tasks or services explicitly mentioned by users, so multiple related domains or tasks need to be performed one by one through a long conversation with many explicit intents.…

Computation and Language · Computer Science 2023-06-07 Hui-Chi Kuo , Yun-Nung Chen

Purpose: In order to produce a surgical gesture recognition system that can support a wide variety of procedures, either a very large annotated dataset must be acquired, or fitted models must generalize to new labels (so called "zero-shot"…

Computer Vision and Pattern Recognition · Computer Science 2024-08-23 Mingxing Rao , Yinhong Qin , Soheil Kolouri , Jie Ying Wu , Daniel Moyer
‹ Prev 1 8 9 10 Next ›