English

Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision

Computer Vision and Pattern Recognition 2026-01-28 v1

Abstract

Despite the significant advancements represented by Vision-Language Models (VLMs), current architectures often exhibit limitations in retaining fine-grained visual information, leading to coarse-grained multimodal comprehension. We attribute this deficiency to a suboptimal training paradigm inherent in prevailing VLMs, which exhibits a text-dominant optimization bias by conceptualizing visual signals merely as passive conditional inputs rather than supervisory targets. To mitigate this, we introduce Youtu-VL, a framework leveraging the Vision-Language Unified Autoregressive Supervision (VLUAS) paradigm, which fundamentally shifts the optimization objective from ``vision-as-input'' to ``vision-as-target.'' By integrating visual tokens directly into the prediction stream, Youtu-VL applies unified autoregressive supervision to both visual details and linguistic content. Furthermore, we extend this paradigm to encompass vision-centric tasks, enabling a standard VLM to perform vision-centric tasks without task-specific additions. Extensive empirical evaluations demonstrate that Youtu-VL achieves competitive performance on both general multimodal tasks and vision-centric tasks, establishing a robust foundation for the development of comprehensive generalist visual agents.

Keywords

Cite

@article{arxiv.2601.19798,
  title  = {Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision},
  author = {Zhixiang Wei and Yi Li and Zhehan Kan and Xinghua Jiang and Zuwei Long and Shifeng Liu and Hongze Shen and Wei Liu and Xiaoyu Tan and Haojia Lin and Yubo Zhu and Qianyu Li and Di Yin and Haoyu Cao and Weibo Gu and Xin Li and Yinsong Liu and Deqiang Jiang and Xing Sun and Yunsheng Wu and Mingkong Tang and Shuangyin Liu and Lexiang Tang and Haodong Lin and Junru Lu and Jiarui Qin and Lingfeng Qiao and Ruizhi Qiao and Bo Ke and Jianfeng He and Ke Li and Yangning Li and Yunhang Shen and Mengdan Zhang and Peixian Chen and Kun Yin and Bing Liu and Yunfei Wu and Huang Chen and Zhongpeng Cai and Xiaotian Li},
  journal= {arXiv preprint arXiv:2601.19798},
  year   = {2026}
}