Wide exploration on robocall surveillance research is hindered due to limited access to public datasets, due to privacy concerns. In this work, we first curate Robo-SAr, a synthetic robocall dataset designed for robocall surveillance research. Robo-SAr comprises of ~200 unwanted and ~1200 legitimate synthetic robocall samples across three realistic adversarial axes: psycholinguistics-manipulated transcripts, emotion-eliciting speech, and cloned voices. We further propose RoboKA, a Kolmogorov-Arnold Network (KAN)-based multimodal fusion framework designed to model structured nonlinear interactions between acoustic and linguistic cues that characterize diverse adversarial robocall strategies. RoboKA first leverages cross-modal contrastive learning to align latent modality representations and feeds the resulting embeddings to a KAN-projection head for final classification. We benchmark RoboKA against strong unimodal and multimodal baselines in both in-domain and out-of-domain setups, finding RoboKA to surpass all baselines in terms of recall and F1-score.
Cite
@article{arxiv.2605.00156,
title = {RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System},
author = {Nitin Choudhury and Nikhil Kumar and Aditya Kumar Sinha and Abhijeet Anand and Hossein Salemi and Orchid Chetia Phukan and Hemant Purohit and Arun Balaji Buduru},
journal= {arXiv preprint arXiv:2605.00156},
year = {2026}
}
Comments
Accepted to the International Conference on Multimedia & Expo (ICME) 2026, 7th International Workshop on Surveillance Data Processing