English

HiLight: Technical Report on the Motern AI Video Language Model

Computer Vision and Pattern Recognition 2024-07-12 v2 Computation and Language Multimedia Image and Video Processing

Abstract

This technical report presents the implementation of a state-of-the-art video encoder for video-text modal alignment and a video conversation framework called HiLight, which features dual visual towers. The work is divided into two main parts: 1.alignment of video and text modalities; 2.convenient and efficient way to interact with users. Our goal is to address the task of video comprehension in the context of billiards. The report includes a discussion of the concepts and the final solution developed during the task's implementation.

Keywords

Cite

@article{arxiv.2407.07325,
  title  = {HiLight: Technical Report on the Motern AI Video Language Model},
  author = {Zhiting Wang and Qiangong Zhou and Kangjie Yang and Zongyang Liu and Xin Mao},
  journal= {arXiv preprint arXiv:2407.07325},
  year   = {2024}
}