English

SituationalLLM: Proactive language models with scene awareness for dynamic, contextual task guidance

Computer Vision and Pattern Recognition 2025-03-05 v3

Abstract

Large language models (LLMs) have achieved remarkable success in text-based tasks but often struggle to provide actionable guidance in real-world physical environments. This is because of their inability to recognize their limited understanding of the user's physical context. We present SituationalLLM, a novel approach that integrates structured scene information into an LLM to deliver proactive, context-aware assistance. By encoding objects, attributes, and relationships in a custom Scene Graph Language, SituationalLLM actively identifies gaps in environmental context and seeks clarifications during user interactions. This behavior emerges from training on the Situational Awareness Database for Instruct-Tuning (SAD-Instruct), which combines diverse, scenario-specific scene graphs with iterative, dialogue-based refinements. Experimental results indicate that SituationalLLM outperforms generic LLM baselines in task specificity, reliability, and adaptability, paving the way for environment-aware AI assistants capable of delivering robust, user-centric guidance under real-world constraints.

Keywords

Cite

@article{arxiv.2406.13302,
  title  = {SituationalLLM: Proactive language models with scene awareness for dynamic, contextual task guidance},
  author = {Muhammad Saif Ullah Khan and Muhammad Zeshan Afzal and Didier Stricker},
  journal= {arXiv preprint arXiv:2406.13302},
  year   = {2025}
}

Comments

Revised Submission to Open Research Europe

R2 v1 2026-06-28T17:11:41.610Z