English

Multimodal Methods for Analyzing Learning and Training Environments: A Systematic Literature Review

Machine Learning 2025-12-19 v2 Multimedia

Abstract

Recent technological advancements in multimodal machine learning--including the rise of large language models (LLMs)--have improved our ability to collect, process, and analyze diverse multimodal data such as speech, video, and eye gaze in learning and training contexts. While prior reviews have addressed individual components of the multimodal pipeline (e.g., conceptual models, data fusion), a comprehensive review of empirical methods in applied multimodal environments remains notably absent. This review addresses that, introducing a taxonomy and framework that capture both established practices and recent innovations driven by LLMs and generative AI. We identify five modality groups: Natural Language, Vision, Physiological Signals, Human-Centered Evidence, and Environment Logs. Our analysis reveals that integrating modalities enables richer insights into learner and trainee behaviors, revealing latent patterns often overlooked by unimodal approaches. However, persistent challenges in multimodal data collection and integration continue to hinder the adoption of these systems in real-time classroom settings.

Keywords

Cite

@article{arxiv.2408.14491,
  title  = {Multimodal Methods for Analyzing Learning and Training Environments: A Systematic Literature Review},
  author = {Clayton Cohn and Eduardo Davalos and Caleb Vatral and Joyce Horn Fonteles and Hanchen David Wang and Austin Coursey and Surya Rayala and Ashwin T S and Meiyi Ma and Gautam Biswas},
  journal= {arXiv preprint arXiv:2408.14491},
  year   = {2025}
}

Comments

Submitted to ACM Computing Surveys. Currently under review

R2 v1 2026-06-28T18:24:19.049Z