English
Related papers

Related papers: Waking Up Blind: Cold-Start Optimization of Superv…

200 papers

Building robust vision systems for high-stakes domains such as remote sensing requires stronger visual reasoning than what single-pass inference typically provides; yet, retraining large models is often computationally expensive and data…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Chung-En Johnny Yu , Brian Jalaian , Nathaniel D. Bastian

Vision-Language Models (VLMs) excel at zero-shot inference but often degrade under test-time domain shifts. For this reason, episodic test-time adaptation strategies have recently emerged as powerful techniques for adapting VLMs to a single…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Konstantinos M. Dafnis , Dimitris N. Metaxas

Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-grained visual reasoning. Recent evidence suggests that this limitation arises not from…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Sophia Sirko-Galouchenko , Monika Wysoczanska , Andrei Bursuc , Nicolas Thome , Spyros Gidaris

Despite significant progress, Vision-Language Models (VLMs) still struggle with complex visual reasoning, where multi-step dependencies cause early errors to cascade through the reasoning chain. Existing post-training paradigms are limited:…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Yanbei Jiang , Chao Lei , Yihao Ding , Krista Ehinger , Jey Han Lau

Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic control, with test-time scaling (TTS) gaining attention to enhance robustness beyond training. However, existing TTS methods for VLAs…

Robotics · Computer Science 2026-02-05 Hyeonbeom Choi , Daechul Ahn , Youhan Lee , Taewook Kang , Seongwon Cho , Jonghyun Choi

Vision-language action (VLA) policies often report strong manipulation benchmark performance with relatively few demonstrations, but it remains unclear whether this reflects robust language-to-object grounding or reliance on…

Robotics · Computer Science 2026-03-02 David Emukpere , Romain Deffayet , Jean-Michel Renders

Vision-language model (VLM) agents increasingly rely on memory-augmented reinforcement learning to reuse experience across long-horizon tasks, yet most existing frameworks store memory as text and depend on proprietary teacher models to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Pan Wang , Yihao Hu , Xiujin Liu , Jingchu Yang , Hang Wang , Zhihao Wen

This work introduces SteerVLM, a lightweight steering module designed to guide Vision-Language Models (VLMs) towards outputs that better adhere to desired instructions. Our approach learns from the latent embeddings of paired prompts…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Anushka Sivakumar , Andrew Zhang , Zaber Hakim , Chris Thomas

With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations…

Computer Vision and Pattern Recognition · Computer Science 2022-05-17 Zirui Wang , Jiahui Yu , Adams Wei Yu , Zihang Dai , Yulia Tsvetkov , Yuan Cao

End-to-end Vision-language Models (VLMs) often answer visual questions by exploiting spurious correlations instead of causal visual evidence, and can become more shortcut-prone when fine-tuned. We introduce VISTA (Visual-Information…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Zhaonan Li , Shijie Lu , Fei Wang , Jacob Dineen , Xiao Ye , Zhikun Xu , Siyi Liu , Young Min Cho , Bangzheng Li , Daniel Chang , Kenny Nguyen , Qizheng Yang , Muhao Chen , Ben Zhou

We introduce VisTA, a new reinforcement learning framework that empowers visual agents to dynamically explore, select, and combine tools from a diverse library based on empirical performance. Existing methods for tool-augmented reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Zeyi Huang , Yuyang Ji , Anirudh Sundara Rajan , Zefan Cai , Wen Xiao , Haohan Wang , Junjie Hu , Yong Jae Lee

Large Language Models (LLMs) are increasingly used for decision-making and planning in autonomous driving, showing promising reasoning capabilities and potential to generalize across diverse traffic situations. However, current LLM-based…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Fabian Schmidt , Noushiq Mohammed Kayilan Abdul Nazar , Markus Enzweiler , Abhinav Valada

Vision-and-Language Navigation (VLN) requires an agent to find a specified spot in an unseen environment by following natural language instructions. Dominant methods based on supervised learning clone expert's behaviours and thus perform…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Hu Wang , Qi Wu , Chunhua Shen

In high-stakes domains, small task-specific vision models are crucial due to their low computational requirements and the availability of numerous methods to explain their results. However, these explanations often reveal that the models do…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Alexander Koebler , Lukas Kuhn , Ingo Thon , Florian Buettner

Perception is crucial for autonomous driving, but single-agent perception is often constrained by sensors' physical limitations, leading to degraded performance under severe occlusion, adverse weather conditions, and when detecting distant…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Xiangbo Gao , Runsheng Xu , Jiachen Li , Ziran Wang , Zhiwen Fan , Zhengzhong Tu

We present SelfPrompt, a novel prompt-tuning approach for vision-language models (VLMs) in a semi-supervised learning setup. Existing methods for tuning VLMs in semi-supervised setups struggle with the negative impact of the miscalibrated…

Computer Vision and Pattern Recognition · Computer Science 2025-01-30 Shuvendu Roy , Ali Etemad

Visual reinforcement learning agents typically face serious performance declines in real-world applications caused by visual distractions. Existing methods rely on fine-tuning the policy's representations with hand-crafted augmentations. In…

Computer Vision and Pattern Recognition · Computer Science 2025-02-17 Xinning Zhou , Chengyang Ying , Yao Feng , Hang Su , Jun Zhu

Recent Vision-Language-Action (VLA) models show strong generalization capabilities, yet they lack introspective mechanisms for anticipating failures and requesting help from a human supervisor. We present \textbf{INSIGHT}, a learning…

Robotics · Computer Science 2026-05-26 Ulas Berk Karli , Ziyao Shangguan , Tesca FItzgerald

State-of-the-art (SOTA) reinforcement learning (RL) methods have enabled vision-language model (VLM) agents to learn from interaction with online environments without human supervision. However, these methods often struggle with learning…

Machine Learning · Computer Science 2025-05-22 Qingyuan Wu , Jianheng Liu , Jianye Hao , Jun Wang , Kun Shao

Reinforcement learning with verifiable outcome rewards (RLVR) has effectively scaled up chain-of-thought (CoT) reasoning in large language models (LLMs). Yet, its efficacy in training vision-language model (VLM) agents for goal-directed…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Tong Wei , Yijun Yang , Junliang Xing , Yuanchun Shi , Zongqing Lu , Deheng Ye
‹ Prev 1 2 3 10 Next ›