Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start from audio alone and actively search for cross-modal evidence remains underexplored. In this paper, we introduce \textbf{Omni-DeepSearch}, a benchmark for audio-driven omni-modal deep search. Given one or more audio clips and a related question, models must infer useful clues from audio, invoke text, image, and video search tools, and perform multi-hop reasoning to produce a short, objective, and verifiable answer. Omni-DeepSearch contains 640 samples across 15 fine-grained categories, covering four retrieval target modalities and four audio content types. A multi-stage filtering pipeline ensures audio dependence, retrieval necessity, visual modality necessity, and answer uniqueness. Experiments on recent closed-source and open-source omni-modal models show that this task remains highly challenging: the strongest evaluated model, Gemini-3-Pro, achieves only 43.44\% average accuracy. Further analyses illustrate key bottlenecks in audio entity inference, query formulation, tool-use reliability, multi-hop retrieval, and cross-modal verification. These results highlight audio-driven omni-modal deep search as an important and underexplored direction for future multimodal agents.
@article{arxiv.2605.08762,
title = {Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search},
author = {Tao Yu and yiming ding and Shenghua Chai and Minghui Zhang and Zhongtian Luo and Xinming Wang and Xinlong Chen and Zhaolu Kang and Junhao Gong and Yuxuan Zhou and Haopeng Jin and Zhiqing Cui and Jiabing Yang and YiFan Zhang and Hongzhu Yi and Zheqi He and Xi Yang and Yan Huang and Liang Wang},
journal= {arXiv preprint arXiv:2605.08762},
year = {2026}
}