English

Kwai Keye-VL 1.5 Technical Report

Computer Vision and Pattern Recognition 2025-09-09 v3

Abstract

In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Models (MLLMs). However, video understanding remains a challenging area due to the dynamic and information-dense nature of videos. Existing models struggle with the trade-off between spatial resolution and temporal coverage when processing video content. We present Keye-VL-1.5, which addresses fundamental challenges in video comprehension through three key innovations. First, we introduce a novel Slow-Fast video encoding strategy that dynamically allocates computational resources based on inter-frame similarity, processing key frames with significant visual changes at higher resolution (Slow pathway) while handling relatively static frames with increased temporal coverage at lower resolution (Fast pathway). Second, we implement a progressive four-stage pre-training methodology that systematically extends the model's context length from 8K to 128K tokens, enabling processing of longer videos and more complex visual content. Third, we develop a comprehensive post-training pipeline focusing on reasoning enhancement and human preference alignment, incorporating a 5-step chain-of-thought data construction process, iterative GSPO-based reinforcement learning with progressive prompt hinting for difficult cases, and alignment training. Through extensive evaluation on public benchmarks and rigorous internal human assessment, Keye-VL-1.5 demonstrates significant improvements over existing models, particularly excelling in video understanding tasks while maintaining competitive performance on general multimodal benchmarks.

Keywords

Cite

@article{arxiv.2509.01563,
  title  = {Kwai Keye-VL 1.5 Technical Report},
  author = {Biao Yang and Bin Wen and Boyang Ding and Changyi Liu and Chenglong Chu and Chengru Song and Chongling Rao and Chuan Yi and Da Li and Dunju Zang and Fan Yang and Guorui Zhou and Guowang Zhang and Han Shen and Hao Peng and Haojie Ding and Hao Wang and Haonan Fan and Hengrui Ju and Jiaming Huang and Jiangxia Cao and Jiankang Chen and Jingyun Hua and Kaibing Chen and Kaiyu Jiang and Kaiyu Tang and Kun Gai and Muhao Wei and Qiang Wang and Ruitao Wang and Sen Na and Shengnan Zhang and Siyang Mao and Sui Huang and Tianke Zhang and Tingting Gao and Wei Chen and Wei Yuan and Xiangyu Wu and Xiao Hu and Xingyu Lu and Yi-Fan Zhang and Yiping Yang and Yulong Chen and Zeyi Lu and Zhenhua Wu and Zhixin Ling and Zhuoran Yang and Ziming Li and Di Xu and Haixuan Gao and Hang Li and Jing Wang and Lejian Ren and Qigen Hu and Qianqian Wang and Shiyao Wang and Xinchen Luo and Yan Li and Yuhang Hu and Zixing Zhang},
  journal= {arXiv preprint arXiv:2509.01563},
  year   = {2025}
}

Comments

Github page: https://github.com/Kwai-Keye/Keye

R2 v1 2026-07-01T05:15:40.654Z