English

Rationalization: A Neural Machine Translation Approach to Generating Natural Language Explanations

Artificial Intelligence 2017-12-20 v2 Computation and Language Human-Computer Interaction Machine Learning

Abstract

We introduce AI rationalization, an approach for generating explanations of autonomous system behavior as if a human had performed the behavior. We describe a rationalization technique that uses neural machine translation to translate internal state-action representations of an autonomous agent into natural language. We evaluate our technique in the Frogger game environment, training an autonomous game playing agent to rationalize its action choices using natural language. A natural language training corpus is collected from human players thinking out loud as they play the game. We motivate the use of rationalization as an approach to explanation generation and show the results of two experiments evaluating the effectiveness of rationalization. Results of these evaluations show that neural machine translation is able to accurately generate rationalizations that describe agent behavior, and that rationalizations are more satisfying to humans than other alternative methods of explanation.

Keywords

Cite

@article{arxiv.1702.07826,
  title  = {Rationalization: A Neural Machine Translation Approach to Generating Natural Language Explanations},
  author = {Upol Ehsan and Brent Harrison and Larry Chan and Mark O. Riedl},
  journal= {arXiv preprint arXiv:1702.07826},
  year   = {2017}
}

Comments

9 pages, 4 figures; added human evaluation section; added author; changed author order-Upol Ehsan and Brent Harrison both contributed equally to this work

R2 v1 2026-06-22T18:28:10.206Z