The Importance of Encoder Choice:A Tabular-Image Study
Abstract
Multimodal learning usually requires a dedicated encoder per modality. When a tabular modality is involved, prior work has been mostly using a \emph{plain MLP} as the encoder. Yet if it were a strong encoder, the tabular domain would not be ``the last unconquered castle for deep learning''. This study evaluates state-of-the-art tabular models as encoders in the image-tabular setting for the first time. An obstacle stands out. In-Context Learning models, among the best performing methods in the tabular domain, require labels to process instances, making it non-trivial to embed training and test instances the same way. We addressed this problem across multiple models of this family. With this study, we would like to highlight the importance of encoder factor in the multimodal learning.
Cite
@article{arxiv.2607.07756,
title = {The Importance of Encoder Choice:A Tabular-Image Study},
author = {Ilia Koloiarov and Diego Coello de Portugal Mecke and Vijaya Krishna Yalavarthi and Tom Hanika and Lars Schmidt-Thieme},
journal= {arXiv preprint arXiv:2607.07756},
year = {2026}
}
Comments
This paper contains color figures. We recommend reading it digitally for the best experience