English

XNOR-FORMER: Learning Accurate Approximations in Long Speech Transformers

Computation and Language 2022-12-21 v2 Artificial Intelligence Sound Audio and Speech Processing

Abstract

Transformers are among the state of the art for many tasks in speech, vision, and natural language processing, among others. Self-attentions, which are crucial contributors to this performance have quadratic computational complexity, which makes training on longer input sequences challenging. Prior work has produced state-of-the-art transformer variants with linear attention, however, current models sacrifice performance to achieve efficient implementations. In this work, we develop a novel linear transformer by examining the properties of the key-query product within self-attentions. Our model outperforms state of the art approaches on speech recognition and speech summarization, resulting in 1 % absolute WER improvement on the Librispeech-100 speech recognition benchmark and a new INTERVIEW speech recognition benchmark, and 5 points on ROUGE for summarization with How2.

Keywords

Cite

@article{arxiv.2210.16643,
  title  = {XNOR-FORMER: Learning Accurate Approximations in Long Speech Transformers},
  author = {Roshan Sharma and Bhiksha Raj},
  journal= {arXiv preprint arXiv:2210.16643},
  year   = {2022}
}

Comments

Under review at ICASSP 2023

R2 v1 2026-06-28T04:46:25.233Z