As Good as It KAN Get: High-Fidelity Audio Representation
Abstract
Implicit neural representations (INR) have gained prominence for efficiently encoding multimedia data, yet their applications in audio signals remain limited. This study introduces the Kolmogorov-Arnold Network (KAN), a novel architecture using learnable activation functions, as an effective INR model for audio representation. KAN demonstrates superior perceptual performance over previous INRs, achieving the lowest Log-SpectralDistance of 1.29 and the highest Perceptual Evaluation of Speech Quality of 3.57 for 1.5 s audio. To extend KAN's utility, we propose FewSound, a hypernetwork-based architecture that enhances INR parameter updates. FewSound outperforms the state-of-the-art HyperSound, with a 33.3% improvement in MSE and 60.87% in SI-SNR. These results show KAN as a robust and adaptable audio representation with the potential for scalability and integration into various hypernetwork frameworks. The source code can be accessed at https://github.com/gmum/fewsound.git.
Keywords
Cite
@article{arxiv.2503.02585,
title = {As Good as It KAN Get: High-Fidelity Audio Representation},
author = {Patryk Marszałek and Maciej Rut and Piotr Kawa and Przemysław Spurek and Piotr Syga},
journal= {arXiv preprint arXiv:2503.02585},
year = {2025}
}
Comments
Accepted to the 34th ACM International Conference on Information and Knowledge Management (CIKM '25)