English

Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts

Computer Vision and Pattern Recognition 2024-03-18 v2 Artificial Intelligence Computation and Language

Abstract

We present a vision-language model whose parameters are jointly trained on all tasks and fully shared among multiple heterogeneous tasks which may interfere with each other, resulting in a single model which we named Musketeer. The integration of knowledge across heterogeneous tasks is enabled by a novel feature called Task Explanation Prompt (TEP). With rich and structured information such as task input/output format, TEP reduces interference among tasks, allowing the model to focus on their shared structure. With a single model, Musketeer achieves results comparable to or better than strong baselines trained on single tasks, almost uniformly across multiple tasks.

Cite

@article{arxiv.2305.07019,
  title  = {Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts},
  author = {Zhaoyang Zhang and Yantao Shen and Kunyu Shi and Zhaowei Cai and Jun Fang and Siqi Deng and Hao Yang and Davide Modolo and Zhuowen Tu and Stefano Soatto},
  journal= {arXiv preprint arXiv:2305.07019},
  year   = {2024}
}
R2 v1 2026-06-28T10:32:19.941Z