English

GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models

Computation and Language 2026-05-29 v1 Artificial Intelligence

Abstract

Deployed language models are evaluated in a non-stationary environment: model versions, retrieval layers, safety systems, and real-world inputs all change over time. Static bias benchmarks remain useful, but they do not show how models frame newly emerging events for different prompted audiences. We introduce GPF-LIVENEWS, a streaming evaluation protocol and benchmark snapshot for auditing group-conditioned framing in open-ended LLM outputs. The protocol expands fresh BBC/Reuters news anchors across 42 identity labels and seven prompt families, then evaluates response bundles using semantic-sensitivity and sentiment-disparity signals. In a pilot over 12 monitoring runs and 23 hosted models, Policy/Action prompts produce the strongest semantic movement, while sentiment variation is flatter across dimensions and prompt families. The released artifact includes article metadata, prompt templates, instantiated prompts, model-output metadata, score tables, documentation, and reproduction scripts. We interpret all scores as observed-window audit signals for human review, not as permanent fairness rankings or direct proof of harmful bias.

Keywords

Cite

@article{arxiv.2605.28848,
  title  = {GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models},
  author = {Mohd Ariful Haque and Fahad Rahman and Kishor Datta Gupta and Roy George},
  journal= {arXiv preprint arXiv:2605.28848},
  year   = {2026}
}
R2 v1 2026-07-22T07:37:52.153Z