We introduce T5Gemma 2, the next generation of the T5Gemma family of lightweight open encoder-decoder models, featuring strong multilingual, multimodal and long-context capabilities. T5Gemma 2 follows the adaptation recipe (via UL2) in T5Gemma -- adapting a pretrained decoder-only model into an encoder-decoder model, and extends it from text-only regime to multimodal based on the Gemma 3 models. We further propose two methods to improve the efficiency: tied word embedding that shares all embeddings across encoder and decoder, and merged attention that unifies decoder self- and cross-attention into a single joint module. Experiments demonstrate the generality of the adaptation strategy over architectures and modalities as well as the unique strength of the encoder-decoder architecture on long context modeling. Similar to T5Gemma, T5Gemma 2 yields comparable or better pretraining performance and significantly improved post-training performance than its Gemma 3 counterpart. We release the pretrained models (270M-270M, 1B-1B and 4B-4B) to the community for future research.
@article{arxiv.2512.14856,
title = {T5Gemma 2: Seeing, Reading, and Understanding Longer},
author = {Biao Zhang and Paul Suganthan and Gaël Liu and Ilya Philippov and Sahil Dua and Ben Hora and Kat Black and Gus Martins and Omar Sanseviero and Shreya Pathak and Cassidy Hardin and Francesco Visin and Jiageng Zhang and Kathleen Kenealy and Qin Yin and Xiaodan Song and Olivier Lacombe and Armand Joulin and Tris Warkentin and Adam Roberts},
journal= {arXiv preprint arXiv:2512.14856},
year = {2025}
}