中文

通过与环境交互解耦独立可控的变异因素

机器学习 2018-02-27 v1 机器学习

摘要

已有假设认为,好的表示应当解耦潜在的解释性变异因素。然而,何种训练框架可能实现这一点仍是一个开放问题。尽管先前大多数工作聚焦于静态设定(例如图像),我们假设若允许学习器与其环境交互,则某些因果因素可被发现。智能体可以尝试不同动作并观察其效果。更具体地,我们假设这些因素中的某些对应于环境中独立可控的方面,即对于环境的每个此类方面,存在一种策略和一个可学习的特征,使得该策略能够引起该特征的变化,同时对解释观测数据中统计变异的其他特征造成最小改变。我们提出了一个特定的目标函数以寻找此类因素,并通过实验验证它确实能够在无任何外在奖励信号的情况下解耦环境的独立可控方面。

关键词

引用

@article{arxiv.1802.09484,
  title  = {Disentangling the independently controllable factors of variation by interacting with the world},
  author = {Valentin Thomas and Emmanuel Bengio and William Fedus and Jules Pondard and Philippe Beaudoin and Hugo Larochelle and Joelle Pineau and Doina Precup and Yoshua Bengio},
  journal= {arXiv preprint arXiv:1802.09484},
  year   = {2018}
}

备注

Presented at NIPS 2017 Learning Disentangling Representations Workshop