English

As Good as It KAN Get: High-Fidelity Audio Representation

Sound 2025-11-04 v3 Computer Vision and Pattern Recognition Audio and Speech Processing

Abstract

Implicit neural representations (INR) have gained prominence for efficiently encoding multimedia data, yet their applications in audio signals remain limited. This study introduces the Kolmogorov-Arnold Network (KAN), a novel architecture using learnable activation functions, as an effective INR model for audio representation. KAN demonstrates superior perceptual performance over previous INRs, achieving the lowest Log-SpectralDistance of 1.29 and the highest Perceptual Evaluation of Speech Quality of 3.57 for 1.5 s audio. To extend KAN's utility, we propose FewSound, a hypernetwork-based architecture that enhances INR parameter updates. FewSound outperforms the state-of-the-art HyperSound, with a 33.3% improvement in MSE and 60.87% in SI-SNR. These results show KAN as a robust and adaptable audio representation with the potential for scalability and integration into various hypernetwork frameworks. The source code can be accessed at https://github.com/gmum/fewsound.git.

Keywords

Cite

@article{arxiv.2503.02585,
  title  = {As Good as It KAN Get: High-Fidelity Audio Representation},
  author = {Patryk Marszałek and Maciej Rut and Piotr Kawa and Przemysław Spurek and Piotr Syga},
  journal= {arXiv preprint arXiv:2503.02585},
  year   = {2025}
}

Comments

Accepted to the 34th ACM International Conference on Information and Knowledge Management (CIKM '25)