边推断边采样:基于朗之万动力学的预测编码
摘要
我们提出了一种用于通用深度生成模型中参数学习的新型算法,该算法建立在计算神经科学的预测编码(PC)框架之上。我们的方法改进了标准 PC 算法,使其性能达到并超过标准变分自编码器(VAE)训练所得性能。通过将高斯噪声注入 PC 推断过程,我们将其重新构想为过阻尼朗之万采样,从而有助于针对紧致证据下界(ELBO)进行优化。我们改进了由此产生的无编码器训练方法,引入编码器网络为我们的朗之万采样提供摊销式热启动,并测试了三种不同的目标函数来实现这一点。最后,为了增加对采样步长的鲁棒性并减少对曲率的敏感性,我们验证了一种轻量且易于计算的前置条件形式,其灵感来自黎曼流形朗之万和 SGD 文献中的自适应优化器。我们通过使用我们的技术训练同类生成模型并与使用标准重参数化技巧 ELBO 训练的模型进行比较来对比 VAE。我们观察到,我们的方法在包括样本质量在内的多项指标上优于或匹配性能,同时在极少部分 SGD 训练迭代次数内收敛。
引用
@article{arxiv.2311.13664,
title = {Sample as You Infer: Predictive Coding With Langevin Dynamics},
author = {Umais Zahid and Qinghai Guo and Zafeirios Fountas},
journal= {arXiv preprint arXiv:2311.13664},
year = {2024}
}
备注
FID values updated to use a fixed 50,000 samples for all experiments - Jeffrey's divergence now consistently best performing. Dynov2 based metrics removed due to inconsistency of results - and since not industry standard. Multiple beta values tested in Fig 4. Theta LR for VAEs; beta and inf LR for LPC now tuned for results. Figure 5B updated; curves now correspond to results in Table 1