中文

Y^2Seq2Seq:通过视图与词序列联合重建与预测实现3D形状与文本跨模态表示学习

计算机视觉与模式识别 2018-11-08 v1

摘要

近期的一种方法采用3D体素来表示3D形状,但由于3D体素的立方复杂度带来的计算成本,该方法局限于低分辨率,因此缺乏精细几何细节。为解决此问题,我们提出基于视图的模型Y^2Seq2Seq,通过视图与词序列的联合重建与预测来学习跨模态表示。具体而言,Y^2Seq2Seq的网络架构通过两个耦合的类似‘Y’的序列到序列(Seq2Seq)结构来桥接两种模态中嵌入的语义。此外,我们新颖的层次化约束通过利用更具判别性的细节信息进一步提升了跨模态表示的可分性。在跨模态检索和3D形状描述生成上的实验结果表明,Y^2Seq2Seq优于最先进的方法。

关键词

引用

@article{arxiv.1811.02745,
  title  = {Y^2Seq2Seq: Cross-Modal Representation Learning for 3D Shape and Text by Joint Reconstruction and Prediction of View and Word Sequences},
  author = {Zhizhong Han and Mingyang Shang and Xiyang Wang and Yu-Shen Liu and Matthias Zwicker},
  journal= {arXiv preprint arXiv:1811.02745},
  year   = {2018}
}

备注

To be pubilished at AAAI 2019