中文

稀疏编码与自编码器

机器学习 2017-10-24 v2 最优化与控制 机器学习

摘要

在“字典学习”中,给定观测 yRny \in \mathbb{R}^ny=Axy = A^*x^*,目标是恢复不相干矩阵 ARn×hA^* \in \mathbb{R}^{n \times h}(通常过完备且假定其列已归一化)以及支撑集大小为 hph^p(其中 0<p<10 <p < 1)的稀疏向量 xRhx^* \in \mathbb{R}^h。在本工作中,我们对自编码器平方损失的梯度下降能否解决字典学习问题进行了严格分析。我们考虑的“自编码器”架构是一个 RnRn\mathbb{R}^n \rightarrow \mathbb{R}^n 映射,包含一个大小为 hh 的单层 ReLU 激活层。在关于 xx^* 非常宽松的分布假设下,我们证明了对于 AA^* 小邻域内的所有点,标准平方损失函数的期望梯度范数在渐近意义上(关于稀疏编码维度)是可以忽略的。合成数据的实验证据支持了这一点。我们还进行了实验以表明 AA^* 是一个局部极小值。在此过程中,我们证明可以设置一层 ReLU 门以自动恢复稀疏编码的支撑集。该性质与损失函数无关。我们认为这可能具有独立的意义。

关键词

引用

@article{arxiv.1708.03735,
  title  = {Sparse Coding and Autoencoders},
  author = {Akshay Rangamani and Anirbit Mukherjee and Amitabh Basu and Tejaswini Ganapathy and Ashish Arora and Sang Chin and Trac D. Tran},
  journal= {arXiv preprint arXiv:1708.03735},
  year   = {2017}
}

备注

In this new version of the paper with a small change in the distributional assumptions we are actually able to prove the asymptotic criticality of a neighbourhood of the ground truth dictionary for even just the standard squared loss of the ReLU autoencoder (unlike the regularized loss in the older version)