English

Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations

Computer Vision and Pattern Recognition 2025-02-17 v2 Artificial Intelligence Machine Learning

Abstract

Imitation learning (IL) enables agents to mimic expert behavior without reward signals but faces challenges in cross-domain scenarios with high-dimensional, noisy, and incomplete visual observations. To address this, we propose Domain-Invariant Per-Frame Feature Extraction for Imitation Learning (DIFF-IL), a novel IL method that extracts domain-invariant features from individual frames and adapts them into sequences to isolate and replicate expert behaviors. We also introduce a frame-wise time labeling technique to segment expert behaviors by timesteps and assign rewards aligned with temporal contexts, enhancing task performance. Experiments across diverse visual environments demonstrate the effectiveness of DIFF-IL in addressing complex visual tasks.

Keywords

Cite

@article{arxiv.2502.02867,
  title  = {Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations},
  author = {Minung Kim and Kawon Lee and Jungmo Kim and Sungho Choi and Seungyul Han},
  journal= {arXiv preprint arXiv:2502.02867},
  year   = {2025}
}

Comments

8 pages main, 19 pages appendix with reference. Submitted to ICML 2025

R2 v1 2026-06-28T21:32:58.277Z