English
Related papers

Related papers: Touch100k: A Large-Scale Touch-Language-Vision Dat…

200 papers

The connection between visual input and tactile sensing is critical for object manipulation tasks such as grasping and pushing. In this work, we introduce the challenging task of estimating a set of tactile physical properties from visual…

Computer Vision and Pattern Recognition · Computer Science 2021-09-22 Matthew Purri , Kristin Dana

Pose estimation is essential for robotic manipulation, particularly when visual perception is occluded during gripper-object interactions. Existing tactile-based methods generally rely on tactile simulation or pre-trained models, which…

Robotics · Computer Science 2026-03-12 Zirui Zhang , Boyang Zhang , Fumin Zhang , Huan Yin

The sense of touch plays a key role in enabling humans to understand and interact with surrounding environments. For robots, tactile sensing is also irreplaceable. While interacting with objects, tactile sensing provides useful information…

Robotics · Computer Science 2021-12-30 Jiaqi Jiang , Shan Luo

Mobile app user interfaces (UIs) are rich with action, text, structure, and image content that can be utilized to learn generic UI representations for tasks like automating user commands, summarizing content, and evaluating the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Andrea Burns , Kate Saenko , Bryan A. Plummer

Robotic manipulation in contact-rich environments remains challenging, particularly when relying on conventional tactile sensors that suffer from limited sensing range, reliability, and cost-effectiveness. In this work, we present LVTG, a…

Robotics · Computer Science 2026-02-04 Yaohua Liu , Binkai Ou , Zicheng Qiu , Ce Hao , Hengjun Zhang

The integration of visual-tactile stimulus is common while humans performing daily tasks. In contrast, using unimodal visual or tactile perception limits the perceivable dimensionality of a subject. However, it remains a challenge to…

Robotics · Computer Science 2019-02-19 Jet-Tsyn Lee , Danushka Bollegala , Shan Luo

Large language models (LLMs) are beginning to automate reward design for dexterous manipulation. However, no prior work has considered tactile sensing, which is known to be critical for human-like dexterity. We present Text2Touch, bringing…

Robotics · Computer Science 2025-09-10 Harrison Field , Max Yang , Yijiong Lin , Efi Psomopoulou , David Barton , Nathan F. Lepora

Visual-language grounding aims to establish semantic correspondences between natural language and visual entities, enabling models to accurately identify and localize target objects based on textual instructions. Existing VLG approaches…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Linfei Li , Lin Zhang , Ying Shen

Textual-based prompt learning methods primarily employ multiple learnable soft prompts and hard class tokens in a cascading manner as text inputs, aiming to align image and text (category) spaces for downstream tasks. However, current…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Zheng Li , Yibing Song , Ming-Ming Cheng , Xiang Li , Jian Yang

We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models. While recent advances in VLA models have introduced robot policies that are both generalizable and…

Multimodal Large Language Models (MLLMs) such as GPT-4V and Gemini Pro face challenges in achieving human-level perception in Visual Question Answering (VQA), particularly in object-oriented perception tasks which demand fine-grained…

Computation and Language · Computer Science 2024-04-09 Songtao Jiang , Yan Zhang , Chenyi Zhou , Yeying Jin , Yang Feng , Jian Wu , Zuozhu Liu

Visual-textual understanding is essential for language-guided robot manipulation. Recent works leverage pre-trained vision-language models to measure the similarity between encoded visual observations and textual instructions, and then…

Robotics · Computer Science 2025-09-30 Chaoran Zhu , Hengyi Wang , Yik Lung Pang , Changjae Oh

In recent years, significant developments have been made in both video retrieval and video moment retrieval tasks, which respectively retrieve complete videos or moments for a given text query. These advancements have greatly improved user…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Ning Han , Yawen Zeng , Shaohua Long , Chengqing Li , Sijie Yang , Dun Tan , Jianfeng Dong , Jingjing Chen

This paper presents a comprehensive survey of vision-language (VL) intelligence from the perspective of time. This survey is inspired by the remarkable progress in both computer vision and natural language processing, and recent trends…

Computer Vision and Pattern Recognition · Computer Science 2022-03-04 Feng Li , Hao Zhang , Yi-Fan Zhang , Shilong Liu , Jian Guo , Lionel M. Ni , PengChuan Zhang , Lei Zhang

This paper aims for the task of text-to-video retrieval, where given a query in the form of a natural-language sentence, it is asked to retrieve videos which are semantically relevant to the given query, from a great number of unlabeled…

Computer Vision and Pattern Recognition · Computer Science 2022-03-04 Jianfeng Dong , Yabing Wang , Xianke Chen , Xiaoye Qu , Xirong Li , Yuan He , Xun Wang

Vision-language models (VLMs) achieve remarkable success in single-image tasks. However, real-world scenarios often involve intricate multi-image inputs, leading to a notable performance decline as models struggle to disentangle critical…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Juntian Zhang , Chuanqi cheng , Yuhan Liu , Wei Liu , Jian Luan , Rui Yan

Visual language tracking (VLT) has emerged as a cutting-edge research area, harnessing linguistic data to enhance algorithms with multi-modal inputs and broadening the scope of traditional single object tracking (SOT) to encompass video…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Xuchen Li , Shiyu Hu , Xiaokun Feng , Dailing Zhang , Meiqi Wu , Jing Zhang , Kaiqi Huang

While vision-language models have advanced significantly, their application in language-conditioned robotic manipulation is still underexplored, especially for contact-rich tasks that extend beyond visually dominant pick-and-place…

Robotics · Computer Science 2025-05-15 Chaofan Zhang , Peng Hao , Xiaoge Cao , Xiaoshuai Hao , Shaowei Cui , Shuo Wang

Text-centric visual question answering (VQA) has made great strides with the development of Multimodal Large Language Models (MLLMs), yet open-source models still fall short of leading models like GPT4V and Gemini, partly due to a lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Jingqun Tang , Chunhui Lin , Zhen Zhao , Shu Wei , Binghong Wu , Qi Liu , Yangfan He , Kuan Lu , Hao Feng , Yang Li , Siqi Wang , Lei Liao , Wei Shi , Yuliang Liu , Hao Liu , Yuan Xie , Xiang Bai , Can Huang

Vision-language tracking has received increasing attention in recent years, as textual information can effectively address the inflexibility and inaccuracy associated with specifying the target object to be tracked. Existing works either…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Xiao Wang , Liye Jin , Xufeng Lou , Shiao Wang , Lan Chen , Bo Jiang , Zhipeng Zhang
‹ Prev 1 4 5 6 7 8 10 Next ›