中文
相关论文

相关论文: Ming-Flash-Omni: A Sparse, Unified Architecture fo…

200 篇论文

Long-form multimodal video understanding requires integrating vision, speech, and ambient audio with coherent long-range reasoning. Existing benchmarks emphasize either temporal length or multimodal richness, but rarely both and while some…

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B…

人工智能 · 计算机科学 2026-05-27 MiniMax , : , Aili Chen , Aonian Li , Baichuan Zhou , Bangwei Gong , Binyang Jiang , Boji Dan , Changqing Yu , Chao Wang , Cheng Ma , Cheng Zhong , Cheng Zhu , Chengjun Xiao , Chengyi Yang , Chengyu Du , Chenyang Zhang , Chi Zhang , Chuangyi Huang , Chunhao Zhang , Chunhui Du , Chunyu Zhao , Congchao Guo , Da Chen , Deming Ding , Dianjun Sun , Dongyu Zhang , Enhui Yang , Fei Yu , Guang Zheng , Guodong Zheng , Guohong Li , Haichao Zhu , Haigang Zhou , Haimo Zhang , Han Ding , Hao Zhang , Haohai Sun , Haolin Lyu , Haonan Lu , Haoyu Wang , Huajie Shi , Huiyang Li , Jiacheng Chen , Jian Zhang , Jiaqi Zhuang , Jiaren Cai , Jiaxin Pan , Jiayao Li , Jiayuan Song , Jichuan Zhang , Jie Wang , Jihao Gu , Jin Zhu , Jingwei Dong , Jingyang Li , Jingyu Zhang , Jingze Zhuang , Jinhao Tian , Jinli Liu , Jinyi Hu , Jun Tao , Jun Zhang , Junbin Ruan , Junhao Xu , Junjie Yan , Junteng Liu , Junxian He , Kang Xu , Ke Ji , Ke Yang , Kecheng Xiao , Keyu Duan , Keyu Li , Le Han , Letian Ruan , Li Yuan , Lianfei Yu , Liheng Feng , Lijie Mo , Lin Li , Lingye Bao , Lingyu Yang , Lingyuan Zhou , Loki , Lu Chen , Lunbin Ceng , Ming Li , Ming Zhong , Mingliang Tao , Mingyuan Chi , Mujie Lin , Nan Hu , Ningxin Chen , Peiyin Zhu , Peng Gao , Pengcheng Gao , Pengfei Li , Penglin Li , Pengyu Zhao , Qibin Ren , Qidi Xu , Qihan Ren , Qile Li , Qin Wang , Quanliang Chen , Qunhong Ceng , Rong Tian , Rui Dong , Ruitao Leng , Ruize Zhang , Shanqi Liu , Shaoyu Chen , Sheng Jia , Shun Yao , Shuoran Zhao , Shuqi Yu , Sichen Li , Sicheng Pan , Songquan Zhu , Tengfei Li , Tian Xie , Tiancheng Qin , Tianrun Liang , Wei Liu , Weiqi Xu , Weitao Li , Weixiang Chen , Weiyu Cheng , Weiyu Zhang , Wenhu Chen , Wenqian Zhao , Xiancai Chen , Xiangjun Song , Xiangyuan Wang , Xiao Luo , Xiao Su , Xiaobo Li , Xiaodong Han , Xiaojie Wu , Xihao Song , Xingyi Han , Xinyu Guan , Xuan Lu , Xun Zou , Xunhao Lai , Xutong Li , Yan Gong , Yang Wang , Yang Xu , Yangsen Wang , Ye Tang , Yicheng Chen , Yinran Qiu , Yiqi Shi , Yiting Guo , Yiwen Huang , Yixuan Wang , Yongyi Hu , Yu Gao , Yu Zhang , Yuanxiang Ying , Yuanzhen Zhang , Yubo Wang , Yuchen Song , Yufeng Yang , Yuhang Meng , Yuhang Miao , Yuhao Li , Yujie Liu , Yulin Hu , Yunan Huang , Yunji Li , Yunyi Huang , Yusen Zhang , Yusu Hong , Yutao Xie , Yutong Zhang , Yuwen Liao , Yuxuan Shi , Yuze Wenren , Zebin Li , Zehan Li , Zejian Luo , Zeyu Jin , Zeyuan Sun , Zhanpeng Zhou , Zhaochen Su , Zhendong Li , Zhengmao Zhu , Zhengyuan Peng , Zhenhua Fan , Zhi Zhang , Zhichao Xu , Zhiheng Lv , Zhikang Xu , Zhitao He , Zhiwei He , Zhongyuan Li , Zibo Gao , Zijia Wu , Zijian Song , Zijian Zhou , Zijun Sun , Zishan Huang , Ziying Chen , Ziyue Ge

We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture processes multimodal image and text inputs, generating…

We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks across four major categories: image restoration, image enhancement, weak-semantic dense…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Yuandong Pu , Le Zhuo , Kaiwen Zhu , Liangbin Xie , Wenlong Zhang , Xiangyu Chen , Peng Gao , Yu Qiao , Chao Dong , Yihao Liu

Multimodal large language models are playing an increasingly significant role in empowering the financial domain, however, the challenges they face, such as multimodal and high-density information and cross-modal multi-hop reasoning, go…

Conventional Vision-Language Models(VLMs) typically utilize a fixed number of vision tokens, regardless of task complexity. This one-size-fits-all strategy introduces notable inefficiencies: using excessive tokens leads to unnecessary…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Junshan Hu , Jialiang Mao , Zhikang Liu , Zhongpu Xia , Peng Jia , Xianpeng Lang

Generative adversarial networks have led to significant advances in cross-modal/domain translation. However, typically these networks are designed for a specific task (e.g., dialogue generation or image synthesis, but not both). We present…

计算机视觉与模式识别 · 计算机科学 2019-07-11 Shuang Ma , Daniel McDuff , Yale Song

The development of large language models (LLMs) has expanded to multi-modal systems capable of processing text, images, and speech within a unified framework. Training these models demands significantly larger datasets and computational…

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Xu Guo , Fulong Ye , Qichao Sun , Liyang Chen , Bingchuan Li , Pengze Zhang , Jiawei Liu , Songtao Zhao , Qian He , Xiangwang Hou

Currently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, streamline deployment, and foster synergistic benefits across different tasks. However, in…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Bin Xia , Yuechen Zhang , Jingyao Li , Chengyao Wang , Yitong Wang , Xinglong Wu , Bei Yu , Jiaya Jia

In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. While the emergence of large language…

Multimodal reasoning has become a cornerstone of modern AI research. Standardized exam questions offer a uniquely rigorous testbed for such reasoning, providing structured visual contexts and verifiable answers. While recent progress has…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Egemen Sert , Şeyda Ertekin

Omni-modal reasoning is essential for intelligent systems to understand and draw inferences from diverse data sources. While existing omni-modal large language models (OLLM) excel at perceiving diverse modalities, they lack the complex…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Yiran Guan , Sifan Tu , Dingkang Liang , Linghao Zhu , Jianzhong Ju , Zhenbo Luo , Jian Luan , Yuliang Liu , Xiang Bai

We introduce Baichuan-Omni-1.5, an omni-modal model that not only has omni-modal understanding capabilities but also provides end-to-end audio generation capabilities. To achieve fluent and high-quality interaction across modalities without…

We introduce OmniFlow, a novel generative model designed for any-to-any generation tasks such as text-to-image, text-to-audio, and audio-to-image synthesis. OmniFlow advances the rectified flow (RF) framework used in text-to-image models to…

In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming…

Real-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction. Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language…

计算与语言 · 计算机科学 2025-05-06 Qingkai Fang , Yan Zhou , Shoutao Guo , Shaolei Zhang , Yang Feng

Vision-Language Models (VLMs) often yield inconsistent descriptions of the same object across viewpoints, hindering the ability of embodied agents to construct consistent semantic representations over time. Previous methods resolved…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Tommaso Galliena , Stefano Rosa , Tommaso Apicella , Pietro Morerio , Alessio Del Bue , Lorenzo Natale

Foundation models have achieved great advances in multi-task learning with a unified interface of unimodal and multimodal tasks. However, the potential of such multi-task learners has not been exploited during transfer learning. In this…

计算机视觉与模式识别 · 计算机科学 2023-05-18 Chengyue Wu , Teng Wang , Yixiao Ge , Zeyu Lu , Ruisong Zhou , Ying Shan , Ping Luo