English

Efficient Text Encoders for Labor Market Analysis

Computation and Language 2025-07-30 v1 Artificial Intelligence

Abstract

Labor market analysis relies on extracting insights from job advertisements, which provide valuable yet unstructured information on job titles and corresponding skill requirements. While state-of-the-art methods for skill extraction achieve strong performance, they depend on large language models (LLMs), which are computationally expensive and slow. In this paper, we propose \textbf{ConTeXT-match}, a novel contrastive learning approach with token-level attention that is well-suited for the extreme multi-label classification task of skill classification. \textbf{ConTeXT-match} significantly improves skill extraction efficiency and performance, achieving state-of-the-art results with a lightweight bi-encoder model. To support robust evaluation, we introduce \textbf{Skill-XL}, a new benchmark with exhaustive, sentence-level skill annotations that explicitly address the redundancy in the large label space. Finally, we present \textbf{JobBERT V2}, an improved job title normalization model that leverages extracted skills to produce high-quality job title representations. Experiments demonstrate that our models are efficient, accurate, and scalable, making them ideal for large-scale, real-time labor market analysis.

Keywords

Cite

@article{arxiv.2505.24640,
  title  = {Efficient Text Encoders for Labor Market Analysis},
  author = {Jens-Joris Decorte and Jeroen Van Hautte and Chris Develder and Thomas Demeester},
  journal= {arXiv preprint arXiv:2505.24640},
  year   = {2025}
}

Comments

This work has been submitted to the IEEE for possible publication

R2 v1 2026-07-01T02:50:43.851Z