Qwen-Audio-3.0-Gen-Preview Technical Report
Abstract
Existing single-domain and multi-task audio systems remain limited in directly organizing speech, music, sound effects, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion Transformer (DiT) and a shared variational autoencoder (VAE) to generate the complete mixed waveform. Prompt enhancement converts free-form requests into structured temporal records that are rendered as textual conditions, while a two-stage data curriculum and semantic conditional views train the proposed model to use these conditions across domains. A shared continuous VAE compresses 48kHz stereo waveforms into 25Hz latent sequences and incorporates semantic supervision, providing one representation for speech, music, sound effects, and their mixtures. On Seed-TTS-Eval, speaker similarity is the proposed model's clearest strength across all three subsets, and on the multi-speaker benchmark, the proposed model shows higher cross-turn consistency than Seed-Audio-1.0 in both languages. On AudioCaps, its advantages are concentrated in evaluations using large audio-language models and AudioBox. Relative to Seed-Audio-1.0, it achieves stronger temporal localization. Using approximately 10% music data of a dedicated in-house model, the proposed model remains close across all seven SongBench components and leads in three while retaining speech and general-audio capabilities. These results demonstrate the potential of unified generation for temporally structured, multi-domain audio.
Cite
@article{arxiv.2607.27011,
title = {Qwen-Audio-3.0-Gen-Preview Technical Report},
author = {Junyu Dai and Xiaoyue Duan and Xinyue Fan and Yihan Feng and Xiangang Li and Yunjia Li and Lejun Min and Yufei Shi and Xingchen Song and Yiran Wang and Cheng Wen and Menglin Wu and Bajian Xiang and Huaicheng Zhang and Han Zhao and Ruichen Zheng},
journal= {arXiv preprint arXiv:2607.27011},
year = {2026}
}