English
Related papers

Related papers: MiMo-V2-Flash Technical Report

200 papers

Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: (1) multi-task pretraining, tasks are co-available at design time where related tasks could borrow representational…

Machine Learning · Computer Science 2026-05-12 Xing Han , Shravan Chaudhari , Tanvi Ranade , Rama Chellappa , Suchi Saria

The Mixture of Experts (MoE) is a widely known neural architecture where an ensemble of specialized sub-models optimizes overall performance with a constant computational cost. However, conventional MoEs pose challenges at scale due to the…

Computation and Language · Computer Science 2023-09-12 Ted Zadouri , Ahmet Üstün , Arash Ahmadian , Beyza Ermiş , Acyr Locatelli , Sara Hooker

Multimodal learning has gained much success in recent years. However, current multimodal fusion methods adopt the attention mechanism of Transformers to implicitly learn the underlying correlation of multimodal features. As a result, the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Thanh-Dat Truong , Christophe Bobda , Nitin Agarwal , Khoa Luu

Mixture-of-Experts (MoE) Large Language Models (LLMs) efficiently scale-up the model while keeping relatively low inference cost. As MoE models only activate part of the experts, related work has proposed expert prediction and caching…

Computation and Language · Computer Science 2025-11-17 Shien Zhu , Samuel Bohl , Robin Oester , Gustavo Alonso

We present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Mixture-of-Modality-Experts (MoME) Transformer, where each…

Computer Vision and Pattern Recognition · Computer Science 2022-05-30 Hangbo Bao , Wenhui Wang , Li Dong , Qiang Liu , Owais Khan Mohammed , Kriti Aggarwal , Subhojit Som , Furu Wei

Text-to-image generation models, especially Multimodal Diffusion Transformers (MMDiT), have shown remarkable progress in generating high-quality images. However, these models often face significant computational bottlenecks, particularly in…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Hanling Zhang , Rundong Su , Zhihang Yuan , Pengtao Chen , Mingzhu Shen Yibo Fan , Shengen Yan , Guohao Dai , Yu Wang

Multimodal large language models (MLLMs) have achieved impressive performance, but high-resolution visual inputs result in long sequences of visual tokens and substantial inference latency. Reducing redundant visual tokens is critical to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Guoyang Xia , Yifeng Ding , Fengfa Li , Lei Ren , Wei Chen , Fangxiang Feng , Xiaojie Wang

We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression Variational Autoencoder, Video-VAE, is designed for video…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Guoqing Ma , Haoyang Huang , Kun Yan , Liangyu Chen , Nan Duan , Shengming Yin , Changyi Wan , Ranchen Ming , Xiaoniu Song , Xing Chen , Yu Zhou , Deshan Sun , Deyu Zhou , Jian Zhou , Kaijun Tan , Kang An , Mei Chen , Wei Ji , Qiling Wu , Wen Sun , Xin Han , Yanan Wei , Zheng Ge , Aojie Li , Bin Wang , Bizhu Huang , Bo Wang , Brian Li , Changxing Miao , Chen Xu , Chenfei Wu , Chenguang Yu , Dapeng Shi , Dingyuan Hu , Enle Liu , Gang Yu , Ge Yang , Guanzhe Huang , Gulin Yan , Haiyang Feng , Hao Nie , Haonan Jia , Hanpeng Hu , Hanqi Chen , Haolong Yan , Heng Wang , Hongcheng Guo , Huilin Xiong , Huixin Xiong , Jiahao Gong , Jianchang Wu , Jiaoren Wu , Jie Wu , Jie Yang , Jiashuai Liu , Jiashuo Li , Jingyang Zhang , Junjing Guo , Junzhe Lin , Kaixiang Li , Lei Liu , Lei Xia , Liang Zhao , Liguo Tan , Liwen Huang , Liying Shi , Ming Li , Mingliang Li , Muhua Cheng , Na Wang , Qiaohui Chen , Qinglin He , Qiuyan Liang , Quan Sun , Ran Sun , Rui Wang , Shaoliang Pang , Shiliang Yang , Sitong Liu , Siqi Liu , Shuli Gao , Tiancheng Cao , Tianyu Wang , Weipeng Ming , Wenqing He , Xu Zhao , Xuelin Zhang , Xianfang Zeng , Xiaojia Liu , Xuan Yang , Yaqi Dai , Yanbo Yu , Yang Li , Yineng Deng , Yingming Wang , Yilei Wang , Yuanwei Lu , Yu Chen , Yu Luo , Yuchu Luo , Yuhe Yin , Yuheng Feng , Yuxiang Yang , Zecheng Tang , Zekai Zhang , Zidong Yang , Binxing Jiao , Jiansheng Chen , Jing Li , Shuchang Zhou , Xiangyu Zhang , Xinhao Zhang , Yibo Zhu , Heung-Yeung Shum , Daxin Jiang

Multi-Head Mixture-of-Experts (MH-MoE) demonstrates superior performance by using the multi-head mechanism to collectively attend to information from various representation spaces within different experts. In this paper, we present a novel…

Computation and Language · Computer Science 2024-12-02 Shaohan Huang , Xun Wu , Shuming Ma , Furu Wei

We introduce Motif-2-12.7B, a new open-weight foundation model that pushes the efficiency frontier of large language models by combining architectural innovation with system-level optimization. Designed for scalable language understanding…

Standard Mixture-of-Experts (MoE) transformers route tokens to expert subnetworks within each layer, but the layer structure itself remains monolithic. We introduce Mixture of Layers (MoL), which replaces full-width transformer blocks…

Machine Learning · Computer Science 2026-05-12 Ivan Ternovtsii , Yurii Bilak

In this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection. We present an improved version of MViT that incorporates decomposed relative…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Yanghao Li , Chao-Yuan Wu , Haoqi Fan , Karttikeya Mangalam , Bo Xiong , Jitendra Malik , Christoph Feichtenhofer

High inter-class similarity, extreme scale variation, and limited computational budgets hinder reliable visual recognition across diverse real-world data. Existing vision-centric and cross-modal approaches often rely on rigid fusion…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Qinghui Chen , Zekai Zhang , Zaigui Zhang , Kai Zhang , Dagang Li , Wenmin Wang , Jinglin Zhang , Cong Liu

Mixture of experts (MoE) is a popular technique to improve capacity of Large Language Models (LLMs) with conditionally-activated parallel experts. However, serving MoE models on memory-constrained devices is challenging due to the large…

Artificial Intelligence · Computer Science 2024-05-30 Rui Kong , Yuanchun Li , Qingtian Feng , Weijun Wang , Xiaozhou Ye , Ye Ouyang , Linghe Kong , Yunxin Liu

Large language models (LLMs) have emerged due to their capability to generate high-quality content across diverse contexts. To reduce their explosively increasing demands for computing resources, a mixture of experts (MoE) has emerged. The…

Hardware Architecture · Computer Science 2025-11-04 Sungmin Yun , Kwanhee Kyung , Juhwan Cho , Jaewan Choi , Jongmin Kim , Byeongho Kim , Sukhan Lee , Kyomin Sohn , Jung Ho Ahn

The attention mechanism has been the core component in modern transformer architectures. However, the computation of standard full attention scales quadratically with the sequence length, serving as a major bottleneck in long-context…

Computation and Language · Computer Science 2026-04-28 Yusheng Zhao , Hourun Li , Bohan Wu , Yichun Yin , Lifeng Shang , Jingyang Yuan , Meng Zhang , Ming Zhang

We introduce LongCat-Flash-Thinking-2601, a 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model with superior agentic reasoning capability. LongCat-Flash-Thinking-2601 achieves state-of-the-art performance among…

Artificial Intelligence · Computer Science 2026-02-03 Meituan LongCat Team , Anchun Gui , Bei Li , Bingyang Tao , Bole Zhou , Borun Chen , Chao Zhang , Chao Zhang , Chen Gao , Chen Zhang , Chengcheng Han , Chenhui Yang , Chuyu Zhang , Cong Chen , Cunguang Wang , Daoru Pan , Defei Bu , Dengchang Zhao , Di Xiu , Dishan Liu , Dongyu Ru , Dunwei Tu , Fan Wu , Fengcheng Yuan , Fengcun Li , Gang Xu , Guanyu Wu , Guoyuan Lin , Haibin Wang , Hansi Yang , Hao Yang , Haonan Yan , Haoxiang Ma , Haoxing Wen , Hongyan Hao , Hongyin Tang , Hongyu Zang , Hongzhi Ni , Hui Su , Jiacheng Zhang , Jiahong Zhou , Jiahuan Li , Jiaming Wang , Jian Yang , Jianfei Zhang , Jianhao Xu , Jianing Wang , Jiapeng Zhu , Jiaqi Sun , Jiarong Shi , Jiarui Zhao , Jingang Wang , Jinluan Yang , Jinrui Ding , Jinwei Xiao , Jiyuan He , Juncan Xu , Kefeng Zhang , Keheng Wang , Li Wei , Lianhui Ma , Lin Qiu , Lingbing Kong , Lingchuan Liu , Linsen Guo , Mengshen Zhu , Mengxia Shen , Mingyang Zhu , Peiguang Li , Peng Pei , Peng Zhao , Pengcheng Jia , Pengtao Zhang , Ping Liu , Qi Gu , Qiong Huang , Qiyuan Duan , Quanchi Weng , Rongxiang Weng , Rongzhi Zhang , Rumei Li , Shanglin Lei , Shengnan An , Shijun Dai , Shizhe Wu , Shuaikang Liu , Shuang Zhou , Shuo Wang , Songyuan Zhao , Tao Liang , Tianhao Hu , Tianze Chen , Wei Liu , Wei Shi , Wei Wang , Weifeng Tang , Wenjie Shi , Wenlong Zhu , Wentao Chen , Wentao Shi , Xi Su , Xiandi Ma , Xiangcheng Liu , Xiangyu Xi , Xiangyuan Liu , Xiangzhou Huang , Xiao Liu , Xiaodong Cai , Xiaolong Chen , Xiaowei Shi , Xiaoyu Li , Xin Chen , Xingchen Liu , Xuan Huang , Xuezhi Cao , Xunliang Cai , Yan Chen , Yang Bai , Yang Liu , Yang Yang , Yang Zheng , Yanyu Chen , Yaoming Wang , Yaoming Zhu , Yaorui Shi , Yaqi Huo , Yerui Sun , Yi Zhang , Yi-Kai Zhang , Yifan Lu , Yifan Zhao , Yihao Chen , Yitao Zhai , Yongjing Yin , Yongwei Zhou , Youshao Xiao , Yu Wang , Yu Yang , Yuchen Xie , Yuchen Yu , Yuchuan Dai , Yue Xu , Yueqing Sun , Yufei Zhang , Yuhuai Wei , Yulei Qian , Yunfan Liang , Yunke Zhao , Yuwei Jiang , Yuxin Bian , Yuxin Chen , Yuxin Liu , Zeyang Yu , Zhao Yang , Zhengsheng Huang , Zhengyu Chen , Zhijian Liu , Zhikang Xia , Zhimin Lin , Zhiyuan Yao , Zhuofan Chen , Zhuowen Han , Zijian Zhang , Ziran Li , Ziwen Wang , Ziyuan Zhuang

We present DeepSeek-Coder-V2, an open-source Mixture-of-Experts (MoE) code language model that achieves performance comparable to GPT4-Turbo in code-specific tasks. Specifically, DeepSeek-Coder-V2 is further pre-trained from an intermediate…

Robust multimodal visual analytics remains challenging when heterogeneous modalities provide complementary but input-dependent evidence for decision-making.Existing multimodal learning methods mainly rely on fixed fusion modules or…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Tianyi Liu , Yiming Li , Wenqian Wang , Jiaojiao Wang , Chen Cai , Yi Wang , Kim-Hui Yap

Computer vision researchers are embracing two promising paradigms: Vision Transformers (ViTs) and Multi-task Learning (MTL), which both show great performance but are computation-intensive, given the quadratic complexity of self-attention…

Hardware Architecture · Computer Science 2023-09-14 Rishov Sarkar , Hanxue Liang , Zhiwen Fan , Zhangyang Wang , Cong Hao