English

Memory-QA: Answering Recall Questions Based on Multimodal Memories

Artificial Intelligence 2025-09-30 v2 Computation and Language Databases

Abstract

We introduce Memory-QA, a novel real-world task that involves answering recall questions about visual content from previously stored multimodal memories. This task poses unique challenges, including the creation of task-oriented memories, the effective utilization of temporal and location information within memories, and the ability to draw upon multiple memories to answer a recall question. To address these challenges, we propose a comprehensive pipeline, Pensieve, integrating memory-specific augmentation, time- and location-aware multi-signal retrieval, and multi-memory QA fine-tuning. We created a multimodal benchmark to illustrate various real challenges in this task, and show the superior performance of Pensieve over state-of-the-art solutions (up to 14% on QA accuracy).

Keywords

Cite

@article{arxiv.2509.18436,
  title  = {Memory-QA: Answering Recall Questions Based on Multimodal Memories},
  author = {Hongda Jiang and Xinyuan Zhang and Siddhant Garg and Rishab Arora and Shiun-Zu Kuo and Jiayang Xu and Ankur Bansal and Christopher Brossman and Yue Liu and Aaron Colak and Ahmed Aly and Anuj Kumar and Xin Luna Dong},
  journal= {arXiv preprint arXiv:2509.18436},
  year   = {2025}
}
R2 v1 2026-07-01T05:51:00.124Z