中文

Valley2:面向可扩展视觉语言模型的多模态设计探索

计算机视觉与模式识别 2025-01-14 v2

摘要

最近,视觉语言模型在诸多任务(如图像描述和视频理解)中取得了显著进展。我们引入 Valley2,一种新型多模态大型语言模型,旨在提升所有领域的性能并拓展电商和短视频场景的实际应用边界。值得注意的是,Valley2 在电商基准测试中实现了超越相似规模开源模型的 SOTA 性能(79.66 vs. 72.76)。此外,Valley2 在参数不足 10B 的模型中名列 OpenCompass 排行榜第二,平均得分高达 67.4。该项目代码与模型权重已开源于 https://github.com/bytedance/Valley。

关键词

引用

@article{arxiv.2501.05901,
  title  = {Valley2: Exploring Multimodal Models with Scalable Vision-Language Design},
  author = {Ziheng Wu and Zhenghao Chen and Ruipu Luo and Can Zhang and Yuan Gao and Zhentao He and Xian Wang and Haoran Lin and Minghui Qiu},
  journal= {arXiv preprint arXiv:2501.05901},
  year   = {2025}
}