English
Related papers

Related papers: Language as an Anchor: Preserving Relative Visual …

200 papers

Visual question answering (VQA) is crucial for promoting surgical education. In practice, the needs of trainees are constantly evolving, such as learning more surgical types, adapting to different robots, and learning new surgical…

Information Retrieval · Computer Science 2024-10-24 Yuyang Du , Kexin Chen , Yue Zhan , Chang Han Low , Tao You , Mobarakol Islam , Ziyu Guo , Yueming Jin , Guangyong Chen , Pheng-Ann Heng

Gaze estimation methods encounter significant performance deterioration when being evaluated across different domains, because of the domain gap between the testing and training data. Existing methods try to solve this issue by reducing the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Guanzhong Zeng , Jingjing Wang , Zefu Xu , Pengwei Yin , Wenqi Ren , Di Xie , Jiang Zhu

We propose LCLA (Language-Conditioned Latent Alignment), a framework for vision-language navigation that learns modular perception-action interfaces by aligning sensory observations to a latent representation of an expert policy. The expert…

Robotics · Computer Science 2026-02-11 Nitesh Subedi , Adam Haroon , Samuel Tetteh , Prajwal Koirala , Cody Fleming , Soumik Sarkar

The wider application of end-to-end learning methods to embodied decision-making domains remains bottlenecked by their reliance on a superabundance of training data representative of the target domain. Meta-reinforcement learning (meta-RL)…

Machine Learning · Computer Science 2024-11-08 Robby Costales , Stefanos Nikolaidis

Vision-Language-Action (VLA) models achieve over 95% success on standard benchmarks. However, through systematic experiments, we find that current state-of-the-art VLA models largely ignore language instructions. Prior work lacks: (1)…

Robotics · Computer Science 2026-03-03 Yuchen Hou , Lin Zhao

Visual and linguistic pre-training aims to learn vision and language representations together, which can be transferred to visual-linguistic downstream tasks. However, there exists semantic confusion between language and vision during the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Shentong Mo , Jingfei Xia , Ihor Markevych

Graph Contrastive Learning frameworks have demonstrated success in generating high-quality node representations. The existing research on efficient data augmentation methods and ideal pretext tasks for graph contrastive learning remains…

Machine Learning · Computer Science 2024-10-22 Zhenyu Lin , Hongzheng Li , Yingxia Shao , Guanhua Ye , Yawen Li , Quanqing Xu

Recent advances in generic large models, such as GPT and DeepSeek, have motivated the introduction of universality to graph pre-training, aiming to learn rich and generalizable knowledge across diverse domains using graph representations to…

Machine Learning · Computer Science 2026-02-27 Lianze Shan , Jitao Zhao , Dongxiao He , Siqi Liu , Jiaxu Cui , Weixiong Zhang

Endeavors have been recently made to transfer knowledge from the labeled pinhole image domain to the unlabeled panoramic image domain via Unsupervised Domain Adaptation (UDA). The aim is to tackle the domain gaps caused by the style…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Xu Zheng , Tianbo Pan , Yunhao Luo , Lin Wang

While current diffusion-based models, typically built on U-Net architectures, have shown promising results on the text-to-motion generation task, they still suffer from semantic misalignment and kinematic artifacts. Through analysis, we…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Haozhe Jia , Wenshuo Chen , Yuqi Lin , Yang Yang , Lei Wang , Mang Ning , Bowen Tian , Songning Lai , Nanqian Jia , Yifan Chen , Yutao Yue

Visual data driven dictionaries have been successfully employed for various object recognition and classification tasks. However, the task becomes more challenging if the training and test data are from contrasting domains. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2014-11-04 Varun Panaganti

The parameter-efficient adaptation of the image-text pretraining model CLIP for video-text retrieval is a prominent area of research. While CLIP is focused on image-level vision-language matching, video-text retrieval demands comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Leqi Shen , Guoqiang Gong , Tianxiang Hao , Tao He , Yifeng Zhang , Pengzhang Liu , Sicheng Zhao , Jungong Han , Guiguang Ding

LLMs typically linearize 2D tables into 1D sequences to fit their autoregressive architecture, which weakens row-column adjacency and other layout cues. In contrast, purely visual encoders can capture spatial cues, yet often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jiancheng Dong , Pengyue Jia , Derong Xu , Jiawei Cheng , Jingyu Peng , Chao Zhang , Bowen Liu , Xin Sun , Lixin Su , Shuaiqiang Wang , Dawei Yin , Xiangyu Zhao

Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric knowledge limits generalization in long-tail scenarios. While Retrieval-Augmented…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Rui Zhao , Haofeng Hu , Zhenhai Gao , Jiaqiao Liu , Gao Fei

Vision-Language Models (VLMs) such as CLIP are trained on large amounts of image-text pairs, resulting in remarkable generalization across several data distributions. However, in several cases, their expensive training and data…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Sravanti Addepalli , Ashish Ramayee Asokan , Lakshay Sharma , R. Venkatesh Babu

Vision-Language-Action (VLA) models aim to control robots for manipulation from visual observations and natural-language instructions. However, existing hierarchical and autoregressive paradigms often introduce architectural overhead,…

Vision Language Navigation (VLN) typically requires agents to navigate to specified objects or remote regions in unknown scenes by obeying linguistic commands. Such tasks require organizing historical visual observations for linguistic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Bolei Chen , Jiaxu Kang , Yifei Wang , Ping Zhong , Qi Wu , Jianxin Wang

Visual grounding (VG) is a challenging task to localize an object in an image based on a textual description. Recent surge in the scale of VG models has substantially improved performance, but also introduced a significant burden on…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Ting Liu , Xuyang Liu , Siteng Huang , Honggang Chen , Quanjun Yin , Long Qin , Donglin Wang , Yue Hu

Vision-language models (VLMs) have advanced rapidly, yet they still struggle with basic spatial reasoning. Despite strong performance on general benchmarks, modern VLMs remain brittle at understanding 2D spatial relationships such as…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Nahid Alam , Leema Krishna Murali , Siddhant Bharadwaj , Patrick Liu , Timothy Chung , Drishti Sharma , Akshata A. , Kranthi Kiran , Wesley Tam , Bala Krishna S Vegesna

Audio-visual joint representation learning under Cross-Modal Generalization (CMG) aims to transfer knowledge from a labeled source modality to an unlabeled target modality through a unified discrete representation space. Existing symmetric…

Machine Learning · Computer Science 2026-02-04 Bixing Wu , Yuhong Zhao , Zongli Ye , Jiachen Lian , Xiangyu Yue , Gopala Anumanchipalli
‹ Prev 1 3 4 5 6 7 10 Next ›