English
Related papers

Related papers: Baichuan-Omni Technical Report

200 papers

The evolution of Omni-Modal Large Language Models~(Omni-LLMs) has revolutionized human--computer interaction, enabling unified audio-visual perception and speech response. However, existing Omni-LLMs struggle with complex real-world…

Sound · Computer Science 2026-03-10 Wenjie Tian , Zhixian Zhao , Jingbin Hu , Huakang Chen , Haohe Liu , Binshen Mu , Lei Xie

The rise in popularity of ChatGPT and GPT-4 has significantly accelerated the development of large models, leading to the creation of numerous impressive large language models(LLMs) and multimodal large language models (MLLMs). These…

Computation and Language · Computer Science 2023-09-18 Conghui He , Zhenjiang Jin , Chao Xu , Jiantao Qiu , Bin Wang , Wei Li , Hang Yan , Jiaqi Wang , Dahua Lin

Recent advances in GPT-4o like multi-modality models have demonstrated remarkable progress for direct speech-to-speech conversation, with real-time speech interaction experience and strong speech understanding ability. However, current…

Sound · Computer Science 2024-12-09 Ze Yuan , Yanqing Liu , Shujie Liu , Sheng Zhao

Recent advances in large language models, particularly following GPT-4o, have sparked increasing interest in developing omni-modal models capable of understanding more modalities. While some open-source alternatives have emerged, there is…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Zuyan Liu , Yuhao Dong , Jiahui Wang , Ziwei Liu , Winston Hu , Jiwen Lu , Yongming Rao

Multimodal Large Language Models (MLLMs) are undergoing rapid progress and represent the frontier of AI development. However, their training and inference efficiency have emerged as a core bottleneck in making MLLMs more accessible and…

Large Multimodal Models (LMMs) have demonstrated exceptional performance across a wide range of domains. This paper explores their potential in pronunciation assessment tasks, with a particular focus on evaluating the capabilities of the…

Sound · Computer Science 2025-03-17 Ke Wang , Lei He , Kun Liu , Yan Deng , Wenning Wei , Sheng Zhao

Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming interaction, yet they still remain far from human-level multimodal interaction. The key…

The reproduction of state-of-the-art multimodal LLM pre-training faces barriers at every stage of the pipeline, including high-quality data filtering, multimodal data mixture strategies, sequence packing techniques, and training frameworks.…

Computation and Language · Computer Science 2025-04-03 Weizhi Wang , Yu Tian , Linjie Yang , Heng Wang , Xifeng Yan

Multimodal Large Language Models (MLLMs) have demonstrated notable capabilities in general visual understanding and reasoning tasks. However, their deployment is hindered by substantial computational costs in both training and inference,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Muyang He , Yexin Liu , Boya Wu , Jianhao Yuan , Yueze Wang , Tiejun Huang , Bo Zhao

As large language models (LLMs) continue to advance, evaluating their comprehensive capabilities becomes significant for their application in various fields. This research study comprehensively evaluates the language, vision, speech, and…

Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for image generation. However, current benchmarks often lack…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Jiayu Wang , Yang Jiao , Yue Yu , Tianwen Qian , Shaoxiang Chen , Jingjing Chen , Yu-Gang Jiang

The rapid development of large language models (LLMs) has spurred extensive research into their domain-specific capabilities, particularly mathematical reasoning. However, most open-source LLMs focus solely on mathematical reasoning,…

Computation and Language · Computer Science 2024-09-04 Shuai Peng , Di Fu , Liangcai Gao , Xiuqin Zhong , Hongguang Fu , Zhi Tang

Recently, Large Language Models (LLMs) have undergone a significant transformation, marked by a rapid rise in both their popularity and capabilities. Leading this evolution are proprietary LLMs like GPT-4 and GPT-o1, which have captured…

We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relative to single-modal counterparts. Qwen3-Omni matches the…

While the recent advances in Multimodal Large Language Models (MLLMs) constitute a significant leap forward in the field, these models are predominantly confined to the realm of input-side multimodal comprehension, lacking the capacity for…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Zhanyu Wang , Longyue Wang , Zhen Zhao , Minghao Wu , Chenyang Lyu , Huayang Li , Deng Cai , Luping Zhou , Shuming Shi , Zhaopeng Tu

Large language models (LLMs) have demonstrated remarkable language abilities. GPT-4, based on advanced LLMs, exhibits extraordinary multimodal capabilities beyond previous visual language models. We attribute this to the use of more…

Computation and Language · Computer Science 2023-05-23 Feilong Chen , Minglun Han , Haozhi Zhao , Qingyang Zhang , Jing Shi , Shuang Xu , Bo Xu

Large multimodal models (LMMs) extend large language models (LLMs) with multi-sensory skills, such as visual understanding, to achieve stronger generic intelligence. In this paper, we analyze the latest model, GPT-4V(ision), to deepen the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Zhengyuan Yang , Linjie Li , Kevin Lin , Jianfeng Wang , Chung-Ching Lin , Zicheng Liu , Lijuan Wang

We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curriculum-inspired progressive training strategy that transitions…

Multimedia · Computer Science 2025-12-01 Meituan LongCat Team , Bairui Wang , Bayan , Bin Xiao , Bo Zhang , Bolin Rong , Borun Chen , Chang Wan , Chao Zhang , Chen Huang , Chen Chen , Chen Chen , Chengxu Yang , Chengzuo Yang , Cong Han , Dandan Peng , Delian Ruan , Detai Xin , Disong Wang , Dongchao Yang , Fanfan Liu , Fengjiao Chen , Fengyu Yang , Gan Dong , Gang Huang , Gang Xu , Guanglu Wan , Guoqiang Tan , Guoqiao Yu , Haibo Qiu , Hao Lu , Hongbo Liu , Hongyu Xiang , Jiaheng Wu , Jian Yang , Jiaxing Liu , Jing Huang , Jingang Wang , Jinrui Ding , Juchao Jiang , Jun Kuang , Jun Wang , Junhui Mei , Ke Ding , Kefeng Zhang , Lei Chen , Liang Shi , Limeng Qiao , Liming Zheng , Lin Ma , Liuyang Guo , Liya Ma , Luying Sun , Man Gao , Mengshen Zhu , Miao Cao , Minliang Lin , Nuo Xu , Peng Shi , Qi Zhang , Qian Fang , Qian Wang , Qian Yang , Quanxiu Wang , Rongxiang Weng , Rongxin Guo , Ruoxuan Liang , Senbin Yang , Shanbo Xu , Shanglin Lei , Shengze Ye , Shimin Chen , Shuaiqi Chen , Shujie Hu , Shuo Li , Siqi Yang , Siyu Xu , Siyu Ren , Song Li , Songxiang Liu , Tianhao Bai , Tianye Dai , Wei Hong , Wei Wang , Weixiao Zhao , Wengang Cao , Wenlong Zhu , Wenlong He , Xi Su , Xi Nan , Xiaohan Zhao , Xiaohao Wang , Xiaoyu Zhao , Xiaoyu Wang , Xiaoyu Li , Xin Pan , Xin Chen , Xiusong Sun , Xu Xiang , Xudong Xing , Xuezhi Cao , Xunliang Cai , Yang Yang , Yanli Tan , Yao Yao , Yerui Sun , Yi Chen , Yifan Lu , Yin Gong , Yining Zhang , Yitian Chen , Yiyang Gan , Yuchen Tang , Yuchen Xie , Yueqian Wang , Yuewen Zheng , Yufei Zhang , Yufeng Zhong , Yulei Qian , Yuqi Peng , Yuqian Li , Yuwei Jiang , Zeyang Hu , Zheng Zhang , Zhengkun Tian , Zhiqing Hong , Zhixiong Zeng , Zhuqi Mi , Ziran Li , Ziwen Wang , Ziyi Zhao , Ziyuan Zhuang , Zizhe Zhao

This paper aims to efficiently enable Large Language Models (LLMs) to use multimodal tools. Advanced proprietary LLMs, such as ChatGPT and GPT-4, have shown great potential for tool usage through sophisticated prompt engineering.…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Rui Yang , Lin Song , Yanwei Li , Sijie Zhao , Yixiao Ge , Xiu Li , Ying Shan

In this report, we introduce InternVL 1.5, an open-source multimodal large language model (MLLM) to bridge the capability gap between open-source and proprietary commercial models in multimodal understanding. We introduce three simple…