English

JFinTEB: Japanese Financial Text Embedding Benchmark

Information Retrieval 2026-04-20 v1 Computation and Language

Abstract

We introduce JFinTEB, the first comprehensive benchmark specifically designed for evaluating Japanese financial text embeddings. Existing embedding benchmarks provide limited coverage of language-specific and domain-specific aspects found in Japanese financial texts. Our benchmark encompasses diverse task categories including retrieval and classification tasks that reflect realistic and well-defined financial text processing scenarios. The retrieval tasks leverage instruction-following datasets and financial text generation queries, while classification tasks cover sentiment analysis, document categorization, and domain-specific classification challenges derived from economic survey data. We conduct extensive evaluations across a wide range of embedding models, including Japanese-specific models of various sizes, multilingual models, and commercial embedding services. We publicly release JFinTEB datasets and evaluation framework at https://github.com/retarfi/JFinTEB to facilitate future research and provide a standardized evaluation protocol for the Japanese financial text mining community. This work addresses a critical gap in Japanese financial text processing resources and establishes a foundation for advancing domain-specific embedding research.

Cite

@article{arxiv.2604.15882,
  title  = {JFinTEB: Japanese Financial Text Embedding Benchmark},
  author = {Masahiro Suzuki and Hiroki Sakaji},
  journal= {arXiv preprint arXiv:2604.15882},
  year   = {2026}
}

Comments

5 pages. Accepted at SIGIR 2026 Resource Track

R2 v1 2026-07-01T12:14:07.705Z