中文

一种基于空间填充曲线的新型音频表示

声音 2022-01-11 v1 音频与语音处理

摘要

由于卷积神经网络(CNNs)已彻底变革图像处理领域,它们已被广泛应用于音频场景。一种常见方法是使用时频分解方法将一维音频信号时间序列转换为二维图像,并且通常会丢弃相位信息。在本文中,我们提出使用空间填充曲线(SFCs)将一维音频波形映射为二维图像。这些映射不压缩输入信号,同时保留其局部结构。此外,这些映射受益于深度学习的进展以及现有大量计算机视觉网络。我们在两个关键词 spotting 问题上测试了八种 SFCs。我们表明,由于 Z 曲线在卷积操作下具有平移等变性,其产生了最佳结果。此外,在多个 CNNs 上,Z 曲线产生了与广泛使用的梅尔频率倒谱系数相当的结果。

关键词

引用

@article{arxiv.2201.02805,
  title  = {A novel audio representation using space filling curves},
  author = {Alessandro Mari and Arash Salarian},
  journal= {arXiv preprint arXiv:2201.02805},
  year   = {2022}
}

备注

2021 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works