English
Related papers

Related papers: The Zamba2 Suite: Technical Report

200 papers

In this report, we introduce Qwen2.5, a comprehensive series of large language models (LLMs) designed to meet diverse needs. Compared to previous iterations, Qwen 2.5 has been significantly improved during both the pre-training and…

The typical Selective State-Space Model (SSM) used in Mamba addresses several limitations of Transformers, such as the quadratic computational complexity with respect to sequence length and the significant memory requirements during…

Computation and Language · Computer Science 2025-10-24 Shengkun Tang , Liqun Ma , Haonan Li , Mingjie Sun , Zhiqiang Shen

Transformers and Mamba, initially invented for natural language processing, have inspired backbone architectures for visual recognition. Recent studies integrated Local Attention Transformers with Mamba to capture both local details and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Meng Lou , Yunxiang Fu , Yizhou Yu

The diffusion model has long been plagued by scalability and quadratic complexity issues, especially within transformer-based structures. In this study, we aim to leverage the long sequence modeling capability of a State-Space Model called…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Vincent Tao Hu , Stefan Andreas Baumann , Ming Gui , Olga Grebenkova , Pingchuan Ma , Johannes Schusterbauer , Björn Ommer

Transformer-based embedding models suffer from quadratic computational and linear memory complexity, limiting their utility for long sequences. We propose recurrent architectures as an efficient alternative, introducing a vertically chunked…

Computation and Language · Computer Science 2026-04-21 Tobias Grantner , Emanuel Sallinger , Martin Flechl

We introduce llama-embed-nemotron-8b, an open-weights text embedding model that achieves state-of-the-art performance on the Multilingual Massive Text Embedding Benchmark (MMTEB) leaderboard as of October 21, 2025. While recent models show…

Computation and Language · Computer Science 2025-11-11 Yauhen Babakhin , Radek Osmulski , Ronay Ak , Gabriel Moreira , Mengyao Xu , Benedikt Schifferer , Bo Liu , Even Oldridge

We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully released artifacts -- model weights, full training data,…

Recently developed large language models (LLMs) such as ChatGPT, Claude, and Llama have demonstrated impressive abilities, and even surpass human-level performance in several tasks. Despite their success, the resource-intensive demands of…

Computation and Language · Computer Science 2024-06-17 Jie Wu , Yufeng Zhu , Lei Shen , Xuqing Lu

We present Hala, a family of Arabic-centric instruction and translation models built with our translate-and-tune pipeline. We first compress a strong AR$\leftrightarrow$EN teacher to FP8 (yielding $\sim$2$\times$ higher throughput with no…

Computation and Language · Computer Science 2025-09-18 Hasan Abed Al Kader Hammoud , Mohammad Zbeeb , Bernard Ghanem

In this report, we introduce our latest translation models, HY-MT1.5-1.8B and HY-MT1.5-7B, a new family of machine translation models developed through a holistic training framework tailored for high-performance translation. Our methodology…

Computation and Language · Computer Science 2026-01-01 Mao Zheng , Zheng Li , Tao Chen , Mingyang Song , Di Wang

The problem of Time-series Forecasting is generally addressed by recurrent, Transformer-based and the recently proposed Mamba-based architectures. However, existing architectures generally process their input at a single temporal scale,…

Machine Learning · Computer Science 2026-03-06 Yusuf Meric Karadag , Ismail Talaz , Ipek Gursel Dino , Sinan Kalkan

As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mamba's long-sequence processing efficiency with Transformer's…

Computation and Language · Computer Science 2025-07-08 Tencent Hunyuan Team , Ao Liu , Botong Zhou , Can Xu , Chayse Zhou , ChenChen Zhang , Chengcheng Xu , Chenhao Wang , Decheng Wu , Dengpeng Wu , Dian Jiao , Dong Du , Dong Wang , Feng Zhang , Fengzong Lian , Guanghui Xu , Guanwei Zhang , Hai Wang , Haipeng Luo , Han Hu , Huilin Xu , Jiajia Wu , Jianchen Zhu , Jianfeng Yan , Jiaqi Zhu , Jihong Zhang , Jinbao Xue , Jun Xia , Junqiang Zheng , Kai Liu , Kai Zhang , Kai Zheng , Kejiao Li , Keyao Wang , Lan Jiang , Lixin Liu , Lulu Wu , Mengyuan Huang , Peijie Yu , Peiqi Wang , Qian Wang , Qianbiao Xiang , Qibin Liu , Qingfeng Sun , Richard Guo , Ruobing Xie , Saiyong Yang , Shaohua Chen , Shihui Hu , Shuai Li , Shuaipeng Li , Shuang Chen , Suncong Zheng , Tao Yang , Tian Zhang , Tinghao Yu , Weidong Han , Weijie Liu , Weijin Zhou , Weikang Wang , Wesleye Chen , Xiao Feng , Xiaoqin Ren , Xingwu Sun , Xiong Kuang , Xuemeng Huang , Xun Cao , Yanfeng Chen , Yang Du , Zhen Yang , Yangyu Tao , Yaping Deng , Yi Shen , Yigeng Hong , Yiqi Chen , Yiqing Huang , Yuchi Deng , Yue Mao , Yulong Wang , Yuyuan Zeng , Zenan Xu , Zhanhui Kang , Zhe Zhao , ZhenXiang Yan , Zheng Fang , Zhichao Hu , Zhongzhi Chen , Zhuoyu Li , Zongwei Li , Alex Yan , Ande Liang , Baitong Liu , Beiping Pan , Bin Xing , Binghong Wu , Bingxin Qu , Bolin Ni , Boyu Wu , Chen Li , Cheng Jiang , Cheng Zhang , Chengjun Liu , Chengxu Yang , Chengzhong Xu , Chiyu Wang , Chong Zha , Daisy Yi , Di Wang , Fanyang Lu , Fei Chen , Feifei Liu , Feng Zheng , Guanghua Yu , Guiyang Li , Guohua Wang , Haisheng Lin , Han Liu , Han Wang , Hao Fei , Hao Lu , Haoqing Jiang , Haoran Sun , Haotian Zhu , Huangjin Dai , Huankui Chen , Huawen Feng , Huihui Cai , Huxin Peng , Jackson Lv , Jiacheng Shi , Jiahao Bu , Jianbo Li , Jianglu Hu , Jiangtao Guan , Jianing Xu , Jianwei Cai , Jiarong Zhang , Jiawei Song , Jie Jiang , Jie Liu , Jieneng Yang , Jihong Zhang , Jin lv , Jing Zhao , Jinjian Li , Jinxing Liu , Jun Zhao , Juntao Guo , Kai Wang , Kan Wu , Lei Fu , Lei He , Lei Wang , Li Liu , Liang Dong , Liya Zhan , Long Cheng , Long Xu , Mao Zheng , Meng Liu , Mengkang Hu , Nanli Chen , Peirui Chen , Peng He , Pengju Pan , Pengzhi Wei , Qi Yang , Qi Yi , Roberts Wang , Rongpeng Chen , Rui Sun , Rui Yang , Ruibin Chen , Ruixu Zhou , Shaofeng Zhang , Sheng Zhang , Shihao Xu , Shuaishuai Chang , Shulin Liu , SiQi Wang , Songjia Feng , Songling Yuan , Tao Zhang , Tianjiao Lang , Tongkai Li , Wei Deng , Wei Li , Weichao Wang , Weigang Zhang , Weixuan Sun , Wen Ouyang , Wenxiang Jiao , Wenzhi Sun , Wenzhuo Jia , Xiang Zhang , Xiangyu He , Xianshun Ren , XiaoYing Zhu , Xiaolong Guo , Xiaoxue Li , Xiaoyu Ma , Xican Lu , Xinhua Feng , Xinting Huang , Xinyu Guan , Xirui Li , Xu Zhang , Xudong Gao , Xun Luo , Xuxiang Qi , Yangkun Chen , Yangyu Tao , Yanling Xiao , Yantao Mai , Yanze Chen , Yao Ding , Yeting Yang , YiFan Song , Yifan Yang , Yijiao Zhu , Yinhe Wu , Yixian Liu , Yong Yang , Yuanjun Cai , Yuanlin Tu , Yue Zhang , Yufei Huang , Yuhang Zhou , Yuhao Jiang , Yuhong Liu , Yuhui Hu , Yujin Lin , Yun Yang , Yunhao Wang , Yusong Zhang , Zekun Wu , Zelong Zhang , Zhan Yu , Zhaoliang Yang , Zhe Zhao , Zheng Li , Zhenyu Huang , Zhiguang Liu , Zhijiang Xu , Zhiqing Kui , Zhiyin Zeng , Zhiyuan Xiong , Zhuo Han , Zifan Wu , Zigang Geng , Zilong Zhao , Ziyan Tang , Ziyuan Zhu , Zonglei Zhu , Zhijiang Xu

With the evolution of large language models, traditional Transformer models become computationally demanding for lengthy sequences due to the quadratic growth in computation with respect to the sequence length. Mamba, emerging as a…

Machine Learning · Computer Science 2024-08-22 Haoran Xu , Ziqian Liu , Rong Fu , Zhongling Su , Zerui Wang , Zheng Cai , Zhilin Pei , Xingcheng Zhang

Sequential recommendation systems aim to predict users' next preferences based on their interaction histories, but existing approaches face critical limitations in efficiency and multi-scale pattern recognition. While Transformer-based…

Information Retrieval · Computer Science 2025-05-08 Qianru Zhang , Liang Qu , Honggang Wen , Dong Huang , Siu-Ming Yiu , Nguyen Quoc Viet Hung , Hongzhi Yin

Large Language Models (LLMs) have played an important role in many fields due to their powerful capabilities.However, their massive number of parameters leads to high deployment requirements and incurs significant inference costs, which…

This paper unveils Dimba, a new text-to-image diffusion model that employs a distinctive hybrid architecture combining Transformer and Mamba elements. Specifically, Dimba sequentially stacked blocks alternate between Transformer and Mamba…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Zhengcong Fei , Mingyuan Fan , Changqian Yu , Debang Li , Youqiang Zhang , Junshi Huang

We present QZhou-Embedding, a general-purpose contextual text embedding model with exceptional text representation capabilities. Built upon the Qwen2.5-7B-Instruct foundation model, we designed a unified multi-task framework comprising…

Computation and Language · Computer Science 2025-09-01 Peng Yu , En Xu , Bin Chen , Haibiao Chen , Yinfei Xu

Large pre-trained models have achieved outstanding results in sequence modeling. The Transformer block and its attention mechanism have been the main drivers of the success of these models. Recently, alternative architectures, such as…

Machine Learning · Computer Science 2025-01-29 J. Pablo Muñoz , Jinjie Yuan , Nilesh Jain

Linear attention transformers have become a strong alternative to softmax attention due to their efficiency. However, linear attention tends to be less expressive and results in reduced accuracy compared to softmax attention. To bridge the…

Machine Learning · Computer Science 2026-05-18 Gabriel Mongaras , Eric C. Larson

Long-term time series forecasting (LTSF) provides longer insights into future trends and patterns. Over the past few years, deep learning models especially Transformers have achieved advanced performance in LTSF tasks. However, LTSF faces…

Machine Learning · Computer Science 2024-06-28 Aobo Liang , Xingguo Jiang , Yan Sun , Xiaohou Shi , Ke Li