English
Related papers

Related papers: Qwen3-Omni Technical Report

200 papers

We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art performance across a diverse range of audio benchmarks, while…

Sound · Computer Science 2025-07-18 Alexander H. Liu , Andy Ehrenberg , Andy Lo , Clément Denoix , Corentin Barreau , Guillaume Lample , Jean-Malo Delignon , Khyathi Raghavi Chandu , Patrick von Platen , Pavankumar Reddy Muddireddy , Sanchit Gandhi , Soham Ghosh , Srijan Mishra , Thomas Foubert , Abhinav Rastogi , Adam Yang , Albert Q. Jiang , Alexandre Sablayrolles , Amélie Héliou , Amélie Martin , Anmol Agarwal , Antoine Roux , Arthur Darcet , Arthur Mensch , Baptiste Bout , Baptiste Rozière , Baudouin De Monicault , Chris Bamford , Christian Wallenwein , Christophe Renaudin , Clémence Lanfranchi , Darius Dabert , Devendra Singh Chaplot , Devon Mizelle , Diego de las Casas , Elliot Chane-Sane , Emilien Fugier , Emma Bou Hanna , Gabrielle Berrada , Gauthier Delerce , Gauthier Guinet , Georgii Novikov , Guillaume Martin , Himanshu Jaju , Jan Ludziejewski , Jason Rute , Jean-Hadrien Chabran , Jessica Chudnovsky , Joachim Studnia , Joep Barmentlo , Jonas Amar , Josselin Somerville Roberts , Julien Denize , Karan Saxena , Karmesh Yadav , Kartik Khandelwal , Kush Jain , Lélio Renard Lavaud , Léonard Blier , Lingxiao Zhao , Louis Martin , Lucile Saulnier , Luyu Gao , Marie Pellat , Mathilde Guillaumin , Mathis Felardos , Matthieu Dinot , Maxime Darrin , Maximilian Augustin , Mickaël Seznec , Neha Gupta , Nikhil Raghuraman , Olivier Duchenne , Patricia Wang , Patryk Saffer , Paul Jacob , Paul Wambergue , Paula Kurylowicz , Philomène Chagniot , Pierre Stock , Pravesh Agrawal , Rémi Delacourt , Romain Sauvestre , Roman Soletskyi , Sagar Vaze , Sandeep Subramanian , Saurabh Garg , Shashwat Dalal , Siddharth Gandhi , Sumukh Aithal , Szymon Antoniak , Teven Le Scao , Thibault Schueller , Thibaut Lavril , Thomas Robert , Thomas Wang , Timothée Lacroix , Tom Bewley , Valeriia Nemychnikova , Victor Paltz , Virgile Richard , Wen-Ding Li , William Marshall , Xuanyu Zhang , Yihan Wan , Yunhao Tang

Audio-language models combine audio encoders with large language models to enable multimodal reasoning, but they also introduce new security vulnerabilities. We propose a universal targeted latent space attack, an encoder-level adversarial…

Sound · Computer Science 2026-01-01 Roee Ziv , Raz Lapid , Moshe Sipper

Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural…

Artificial Intelligence · Computer Science 2025-10-23 Brandon James Carone , Iran R. Roman , Pablo Ripollés

We present AM-Thinking-v1, a 32B dense language model that advances the frontier of reasoning, embodying the collaborative spirit of open-source innovation. Outperforming DeepSeek-R1 and rivaling leading Mixture-of-Experts (MoE) models like…

Computation and Language · Computer Science 2025-05-27 Yunjie Ji , Xiaoyu Tian , Sitong Zhao , Haotian Wang , Shuaiting Chen , Yiping Peng , Han Zhao , Xiangang Li

We propose Easy End-to-End Diffusion-based Text to Speech, a simple and efficient end-to-end text-to-speech model based on diffusion. E3 TTS directly takes plain text as input and generates an audio waveform through an iterative refinement…

Sound · Computer Science 2023-11-03 Yuan Gao , Nobuyuki Morioka , Yu Zhang , Nanxin Chen

Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains, mathematical reasoning serves as a representative benchmark…

We present Qwen-Image-VAE-2.0, a suite of high-compression Variational Autoencoders (VAEs) that achieve significant advances in both reconstruction fidelity and diffusability. To address the reconstruction bottlenecks of high compression,…

We present Qwen3-Coder-Next, an open-weight language model specialized for coding agents. Qwen3-Coder-Next is an 80-billion-parameter model that activates only 3 billion parameters during inference, enabling strong coding capability with…

Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over omni-modal (visual, audio, and textual) signals in long and…

Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on edge devices. We present Mobile-O, a compact…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Abdelrahman Shaker , Ahmed Heakl , Jaseel Muhammad , Ritesh Thawkar , Omkar Thawakar , Senmao Li , Hisham Cholakkal , Ian Reid , Eric P. Xing , Salman Khan , Fahad Shahbaz Khan

We present a multi-expert system for creating Non-Player Characters (NPCs) capable of both natural dialogue and contextual action execution in interactive environments. Using Qwen3 as the base model and Low-Rank Adaptation (LoRA) adapters,…

Computation and Language · Computer Science 2025-11-04 Mahammad Nuriyev

Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak their answers. To study this, we introduce OmniACBench, a benchmark for evaluating…

Computation and Language · Computer Science 2026-03-26 Seunghee Kim , Bumkyu Park , Kyudan Jung , Joosung Lee , Soyoon Kim , Jeonghoon Kim , Taeuk Kim , Hwiyeol Jo

With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by leveraging the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Rongtao Xu , Han Gao , Mingming Yu , Dong An , Shunpeng Chen , Changwei Wang , Li Guo , Xiaodan Liang , Shibiao Xu

Voice agents, artificial intelligence systems that conduct spoken conversations to complete tasks, are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses two core evaluation challenges:…

We introduce Full-Duplex-Bench-v3 (FDB-v3), a benchmark for evaluating spoken language models under naturalistic speech conditions and multi-step tool use. Unlike prior work, our dataset consists entirely of real human audio annotated for…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-07 Guan-Ting Lin , Chen Chen , Zhehuai Chen , Hung-yi Lee

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Gaojie Lin , Jianwen Jiang , Jiaqi Yang , Zerong Zheng , Chao Liang

Omnimodal understanding entails a massive, highly redundant search space of cross-modal interactions, demanding focused and deliberative reasoning. Current reasoning paradigms rely on either sequential step-by-step generation or parallel…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Zhicheng Zhang , Wentao Gu , Weicheng Wang , Yongjie Zhu , Wenyu Qin , Meng Wang , Pengfei Wan , Jufeng Yang

Despite recent advances in speech-to-speech translation (S2ST), it remains difficult to achieve both high translation accuracy and practical flexibility. In this paper, we present S2ST-Omni, a compositional S2ST framework that integrates a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-07 Yu Pan , Xiongfei Wu , Yuguang Yang , Jixun Yao , Lei Ma , Jianjun Zhao

We present CM3Leon (pronounced "Chameleon"), a retrieval-augmented, token-based, decoder-only multi-modal language model capable of generating and infilling both text and images. CM3Leon uses the CM3 multi-modal architecture but…

Classroom discourse is an essential vehicle through which teaching and learning take place. Assessing different characteristics of discursive practices and linking them to student learning achievement enhances the understanding of teaching…

Computers and Society · Computer Science 2025-05-14 Ruikun Hou , Babette Bühler , Tim Fütterer , Efe Bozkir , Peter Gerjets , Ulrich Trautwein , Enkelejda Kasneci
‹ Prev 1 8 9 10 Next ›