English
Related papers

Related papers: Qwen3-Omni Technical Report

200 papers

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatenate representation of…

Artificial Intelligence · Computer Science 2025-06-24 Shaolei Zhang , Shoutao Guo , Qingkai Fang , Yan Zhou , Yang Feng

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its…

Machine Learning · Computer Science 2026-05-12 NVIDIA , : , Amala Sanjay Deshmukh , Kateryna Chumachenko , Tuomas Rintamaki , Matthieu Le , Tyler Poon , Danial Mohseni Taheri , Ilia Karmanov , Guilin Liu , Jarno Seppanen , Arushi Goel , Mike Ranzinger , Greg Heinrich , Guo Chen , Lukas Voegtle , Philipp Fischer , Timo Roman , Karan Sapra , Collin McCarthy , Shaokun Zhang , Fuxiao Liu , Hanrong Ye , Yi Dong , Mingjie Liu , Yifan Peng , Piotr Zelasko , Zhehuai Chen , Nithin Rao Koluguri , Nune Tadevosyan , Lilit Grigoryan , Ehsan Hosseini Asl , Pritam Biswas , Leili Tavabi , Yuanhang Su , Zhiding Yu , Peter Jin , Alexandre Milesi , Netanel Haber , Yao Xu , Sarah Amiraslani , Nabin Mulepati , Eric Tramel , Jaehun Jung , Ximing Lu , Brandon Cui , Jin Xu , Zhiqi Li , Shihao Wang , Yuanguo Kuang , Shaokun Zhang , Huck Yang , Boyi Li , Hongxu Yin , Song Han , Bilal Kartal , Pavlo Molchanov , Adi Renduchintala , Charles Wang , David Mosallanezhad , Soumye Singhal , Luis Vega , Katherine Cheung , Sreyan Ghosh , Yian Zhang , Alexander Bukharin , Venkat Srinivasan , Johnny Greco , Andre Manoel , Maarten Van Segbroeck , Suseella Panguliri , Rohit Watve , Divyanshu Kakwani , Shubham Pachori , Jeffrey Glick , Radha Sri-Tharan , Aileen Zaman , Khanh Nguyen , Shi Chen , Jiaheng Fang , Qing Miao , Wenfei Zhou , Yu Wang , Zaid Pervaiz Bhat , Varun Praveen , Arihant Jain , Ramanathan Arunachalam , Tomasz Kornuta , Ashton Sharabiani , Amy Shen , Wei Huang , Yi-Fu Wu , Ali Roshan Ghias , Huiying Li , Brian Yu , Nima Tajbakhsh , Chen Cui , Wenwen Gao , Li Ding , Terry Kong , Manoj Kilaru , Anahita Bhiwandiwalla , Marek Wawrzos , Daniel Korzekwa , Pablo Ribalta , Grzegorz Chlebus , Besmira Nushi , Ewa Dobrowolska , Maciej Jakub Mikulski , Kunal Dhawan , Steve Huang , Jagadeesh Balam , Yongqiang Wang , Nikolay Karpov , Valentin Mendelev , George Zelenfroynd , Meline Mkrtchyan , Qing Miao , Omri Almog , Bhavesh Pawar , Rameshwar Shivbhakta , Sudeep Sabnis , Ashrton Sharabiani , Negar Habibi , Geethapriya Venkataramani , Pamela Peng , Prerit Rodney , Serge Panev , Richard Mazzarese , Nicky Liu , Michael Fukuyama , Andrii Skliar , Roger Waleffe , Duncan Riach , Yunheng Zou , Jian Hu , Hao Zhang , Binfeng Xu , Yuhao Yang , Zuhair Ahmed , Alexandre Milesi , Carlo del Mundo , Chad Voegele , Zhiyu Cheng , Nave Assaf , Andrii Skliar , Daniel Afrimi , Natan Bagrov , Ran Zilberstein , Ofri Masad , Eugene Khvedchenia , Natan Bagrov , Borys Tymchenko , Tomer Asida , Daniel Afrimi , Parth Mannan , Victor Cui , Michael Evans , Katherine Luna , Jie Lou , Pinky Xu , Guyue Huang , Negar Habibi , Michael Boone , Pradeep Thalasta , Adeola Adesoba , Dina Yared , Christopher Parisien , Leon Derczynski , Shaona Ghosh , Wes Feely , Micah Schaffer , Radha Sri-Tharan , Jeffrey Glick , Barnaby Simkin , George Zelenfroynd , Tomasz Grzegorzek , Rishabh Garg , Aastha Jhunjhunwala , Sergei Kolchenko , Farzan Memarian , Haran Kumar , Shiv Kumar , Isabel Hulseman , Anjali Shah , Kari Briski , Padmavathy Subramanian , Joey Conway , Udi Karpas , Jane Polak Scowcroft , Annie Surla , Shilpa Ammireddy , Ellie Evans , Jesse Oliver , Tom Balough , Chia-Chih Chen , Sandip Bhaskar , Alejandra Rico , Bardiya Sadeghi , Seph Mard , Katherine Cheung , Meredith Price , Laya Sleiman , Saori Kaji , Wesley Helmholz , Wendy Quan , Michael Lightstone , Jonathan Cohen , Jian Zhang , Oleksii Kuchaiev , Boris Ginsburg , Jan Kautz , Eileen Long , Mohammad Shoeybi , Mostofa Patwary , Oluwatobi Olabiyi , Andrew Tao , Bryan Catanzaro , Udi Karpas

Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasoning requires fine-grained evidence from both modalities. A central limitation is that…

The evolution of Omni-Modal Large Language Models~(Omni-LLMs) has revolutionized human--computer interaction, enabling unified audio-visual perception and speech response. However, existing Omni-LLMs struggle with complex real-world…

Sound · Computer Science 2026-03-10 Wenjie Tian , Zhixian Zhao , Jingbin Hu , Huakang Chen , Haohe Liu , Binshen Mu , Lei Xie

We present M2-omni, a cutting-edge, open-source omni-MLLM that achieves competitive performance to GPT-4o. M2-omni employs a unified multimodal sequence modeling framework, which empowers Large Language Models(LLMs) to acquire comprehensive…

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually solvable or only…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Le Thien Phuc Nguyen , Zhuoran Yu , Samuel Low Yu Hang , Subin An , Jeongik Lee , Yohan Ban , SeungEun Chung , Thanh-Huy Nguyen , JuWan Maeng , Soochahn Lee , Yong Jae Lee

We present Uni-MoE 2.0 from the Lychee family. As a fully open-source omnimodal large model (OLM), it substantially advances Lychee's Uni-MoE series in language-centric multimodal understanding, reasoning, and generating. Based on the dense…

Computation and Language · Computer Science 2025-11-25 Yunxin Li , Xinyu Chen , Shenyuan Jiang , Haoyuan Shi , Zhenyu Liu , Xuanyu Zhang , Nanhao Deng , Zhenran Xu , Yicheng Ma , Meishan Zhang , Baotian Hu , Min Zhang

Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily focus on short audio and video clips ranging from 10 seconds…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Keda Tao , Yuhua Zheng , Jia Xu , Wenjie Du , Kele Shao , Hesong Wang , Xueyi Chen , Xin Jin , Junhan Zhu , Bohan Yu , Weiqiang Wang , Jian Liu , Can Qin , Yulun Zhang , Ming-Hsuan Yang , Huan Wang

MiniMind-O is an open 0.1B-scale omni model built on the MiniMind language model. It accepts text, speech, and image inputs, and returns both text and streaming speech. The release includes model code, checkpoints, and the main Parquet…

Sound · Computer Science 2026-05-06 Jingyao Gong

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Zhongjian Wang , Peng Zhang , Jinwei Qi , Guangyuan Wang , Chaonan Ji , Sheng Xu , Bang Zhang , Liefeng Bo

Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio-visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective…

Computation and Language · Computer Science 2026-01-21 Qian Chen , Jinlan Fu , Changsong Li , See-Kiong Ng , Xipeng Qiu

Recently, the powerful text-to-image capabilities of ChatGPT-4o have led to growing appreciation for native multimodal large language models. However, its multimodal capabilities remain confined to images and text. Yet beyond images, the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Junliang Ye , Zhengyi Wang , Ruowen Zhao , Shenghao Xie , Jun Zhu

Recent advancements in multimodal large language models (MLLMs) have aimed to integrate and interpret data across diverse modalities. However, the capacity of these models to concurrently process and reason about multiple modalities remains…

Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study…

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, audio-dense} design --…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Detao Bai , Shimin Yao , Weixuan Chen , Chengen Lai , Yuanming Li , Zhiheng Ma , Xihan Wei

In this work, we introduce the Qwen3 Embedding series, a significant advancement over its predecessor, the GTE-Qwen series, in text embedding and reranking capabilities, built upon the Qwen3 foundation models. Leveraging the Qwen3 LLMs'…

Computation and Language · Computer Science 2025-06-12 Yanzhao Zhang , Mingxin Li , Dingkun Long , Xin Zhang , Huan Lin , Baosong Yang , Pengjun Xie , An Yang , Dayiheng Liu , Junyang Lin , Fei Huang , Jingren Zhou

Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-modal large language models. Existing benchmarks fall short in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Ruixiang Zhao , Jie Yang , Zijie Xin , Tianyi Wang , Fengyun Rao , Jing LYU , Xirong Li

We present BlueLM-2.5-3B, a compact and unified dense Multimodal Large Language Model (MLLM) designed for efficient edge-device deployment, offering strong general-purpose and reasoning capabilities. To the best of our knowledge, this is…

Recent advances in GPT-4o like multi-modality models have demonstrated remarkable progress for direct speech-to-speech conversation, with real-time speech interaction experience and strong speech understanding ability. However, current…

Sound · Computer Science 2024-12-09 Ze Yuan , Yanqing Liu , Shujie Liu , Sheng Zhao

The rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despite their potential,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Yuxuan Wang , Yueqian Wang , Bo Chen , Tong Wu , Dongyan Zhao , Zilong Zheng