English
Related papers

Related papers: Jamba: A Hybrid Transformer-Mamba Language Model

200 papers

Image generation models have encountered challenges related to scalability and quadratic complexity, primarily due to the reliance on Transformer-based backbones. In this study, we introduce MaskMamba, a novel hybrid model that combines…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Wenchao Chen , Liqiang Niu , Ziyao Lu , Fandong Meng , Jie Zhou

With the evolution of large language models, traditional Transformer models become computationally demanding for lengthy sequences due to the quadratic growth in computation with respect to the sequence length. Mamba, emerging as a…

Machine Learning · Computer Science 2024-08-22 Haoran Xu , Ziqian Liu , Rong Fu , Zhongling Su , Zerui Wang , Zheng Cai , Zhilin Pei , Xingcheng Zhang

Linear State Space Models (SSMs) offer remarkable performance gains in efficient sequence modeling, with constant inference-time computation and memory complexity. Recent advances, such as Mamba, further enhance SSMs with input-dependent…

Machine Learning · Computer Science 2025-06-24 Zheng Zhan , Liliang Ren , Shuohang Wang , Liyuan Liu , Yang Liu , Yeyun Gong , Yanzhi Wang , Yelong Shen

Mamba, a special case of the State Space Model, is gaining popularity as an alternative to template-based deep learning approaches in medical image analysis. While transformers are powerful architectures, they have drawbacks, including…

Transformer-based models have become increasingly popular and have impacted speech-processing research owing to their exceptional performance in sequence modeling. Recently, a promising model architecture, Mamba, has emerged as a potential…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-27 Wen-Yuan Ting , Wenze Ren , Rong Chao , Hsin-Yi Lin , Yu Tsao , Fan-Gang Zeng

We introduce a novel deep learning method for decoding error correction codes based on the Mamba architecture, enhanced with Transformer layers. Our approach proposes a hybrid decoder that leverages Mamba's efficient sequential modeling…

Information Theory · Computer Science 2025-05-26 Shy-el Cohen , Yoni Choukroun , Eliya Nachmani

As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mamba's long-sequence processing efficiency with Transformer's…

Computation and Language · Computer Science 2025-07-08 Tencent Hunyuan Team , Ao Liu , Botong Zhou , Can Xu , Chayse Zhou , ChenChen Zhang , Chengcheng Xu , Chenhao Wang , Decheng Wu , Dengpeng Wu , Dian Jiao , Dong Du , Dong Wang , Feng Zhang , Fengzong Lian , Guanghui Xu , Guanwei Zhang , Hai Wang , Haipeng Luo , Han Hu , Huilin Xu , Jiajia Wu , Jianchen Zhu , Jianfeng Yan , Jiaqi Zhu , Jihong Zhang , Jinbao Xue , Jun Xia , Junqiang Zheng , Kai Liu , Kai Zhang , Kai Zheng , Kejiao Li , Keyao Wang , Lan Jiang , Lixin Liu , Lulu Wu , Mengyuan Huang , Peijie Yu , Peiqi Wang , Qian Wang , Qianbiao Xiang , Qibin Liu , Qingfeng Sun , Richard Guo , Ruobing Xie , Saiyong Yang , Shaohua Chen , Shihui Hu , Shuai Li , Shuaipeng Li , Shuang Chen , Suncong Zheng , Tao Yang , Tian Zhang , Tinghao Yu , Weidong Han , Weijie Liu , Weijin Zhou , Weikang Wang , Wesleye Chen , Xiao Feng , Xiaoqin Ren , Xingwu Sun , Xiong Kuang , Xuemeng Huang , Xun Cao , Yanfeng Chen , Yang Du , Zhen Yang , Yangyu Tao , Yaping Deng , Yi Shen , Yigeng Hong , Yiqi Chen , Yiqing Huang , Yuchi Deng , Yue Mao , Yulong Wang , Yuyuan Zeng , Zenan Xu , Zhanhui Kang , Zhe Zhao , ZhenXiang Yan , Zheng Fang , Zhichao Hu , Zhongzhi Chen , Zhuoyu Li , Zongwei Li , Alex Yan , Ande Liang , Baitong Liu , Beiping Pan , Bin Xing , Binghong Wu , Bingxin Qu , Bolin Ni , Boyu Wu , Chen Li , Cheng Jiang , Cheng Zhang , Chengjun Liu , Chengxu Yang , Chengzhong Xu , Chiyu Wang , Chong Zha , Daisy Yi , Di Wang , Fanyang Lu , Fei Chen , Feifei Liu , Feng Zheng , Guanghua Yu , Guiyang Li , Guohua Wang , Haisheng Lin , Han Liu , Han Wang , Hao Fei , Hao Lu , Haoqing Jiang , Haoran Sun , Haotian Zhu , Huangjin Dai , Huankui Chen , Huawen Feng , Huihui Cai , Huxin Peng , Jackson Lv , Jiacheng Shi , Jiahao Bu , Jianbo Li , Jianglu Hu , Jiangtao Guan , Jianing Xu , Jianwei Cai , Jiarong Zhang , Jiawei Song , Jie Jiang , Jie Liu , Jieneng Yang , Jihong Zhang , Jin lv , Jing Zhao , Jinjian Li , Jinxing Liu , Jun Zhao , Juntao Guo , Kai Wang , Kan Wu , Lei Fu , Lei He , Lei Wang , Li Liu , Liang Dong , Liya Zhan , Long Cheng , Long Xu , Mao Zheng , Meng Liu , Mengkang Hu , Nanli Chen , Peirui Chen , Peng He , Pengju Pan , Pengzhi Wei , Qi Yang , Qi Yi , Roberts Wang , Rongpeng Chen , Rui Sun , Rui Yang , Ruibin Chen , Ruixu Zhou , Shaofeng Zhang , Sheng Zhang , Shihao Xu , Shuaishuai Chang , Shulin Liu , SiQi Wang , Songjia Feng , Songling Yuan , Tao Zhang , Tianjiao Lang , Tongkai Li , Wei Deng , Wei Li , Weichao Wang , Weigang Zhang , Weixuan Sun , Wen Ouyang , Wenxiang Jiao , Wenzhi Sun , Wenzhuo Jia , Xiang Zhang , Xiangyu He , Xianshun Ren , XiaoYing Zhu , Xiaolong Guo , Xiaoxue Li , Xiaoyu Ma , Xican Lu , Xinhua Feng , Xinting Huang , Xinyu Guan , Xirui Li , Xu Zhang , Xudong Gao , Xun Luo , Xuxiang Qi , Yangkun Chen , Yangyu Tao , Yanling Xiao , Yantao Mai , Yanze Chen , Yao Ding , Yeting Yang , YiFan Song , Yifan Yang , Yijiao Zhu , Yinhe Wu , Yixian Liu , Yong Yang , Yuanjun Cai , Yuanlin Tu , Yue Zhang , Yufei Huang , Yuhang Zhou , Yuhao Jiang , Yuhong Liu , Yuhui Hu , Yujin Lin , Yun Yang , Yunhao Wang , Yusong Zhang , Zekun Wu , Zelong Zhang , Zhan Yu , Zhaoliang Yang , Zhe Zhao , Zheng Li , Zhenyu Huang , Zhiguang Liu , Zhijiang Xu , Zhiqing Kui , Zhiyin Zeng , Zhiyuan Xiong , Zhuo Han , Zifan Wu , Zigang Geng , Zilong Zhao , Ziyan Tang , Ziyuan Zhu , Zonglei Zhu , Zhijiang Xu

We propose TRAMBA, a hybrid transformer and Mamba architecture for acoustic and bone conduction speech enhancement, suitable for mobile and wearable platforms. Bone conduction speech enhancement has been impractical to adopt in mobile and…

Sound · Computer Science 2024-05-30 Yueyuan Sui , Minghui Zhao , Junxi Xia , Xiaofan Jiang , Stephen Xia

We propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates transformer attention mechanisms with state space models (SSMs) for enhanced efficiency. Attention heads provide…

Transformers have become increasingly popular for image super-resolution (SR) tasks due to their strong global context modeling capabilities. However, their quadratic computational complexity necessitates the use of window-based attention…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Aman Urumbekov , Zheng Chen

Balancing fine-grained local modeling with long-range dependency capture under computational constraints remains a central challenge in sequence modeling. While Transformers provide strong token mixing, they suffer from quadratic…

Machine Learning · Computer Science 2026-03-20 Youjin Wang , Jiaqiao Zhao , Rong Fu , Run Zhou , Ruizhe Zhang , Jiani Liang , Suisuai Cao , Feng Zhou

Transformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with quadratic complexity is inefficient in computation and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-02 Xilin Jiang , Cong Han , Nima Mesgarani

Multimodal Large Language Models (MLLMs) have attracted much attention for their multifunctionality. However, traditional Transformer architectures incur significant overhead due to their secondary computational complexity. To address this…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Wenjun Huang , Jiakai Pan , Jiahao Tang , Yanyu Ding , Yifei Xing , Yuhe Wang , Zhengzhuo Wang , Jianguo Hu

State-of-the-art transformer-based large multimodal models (LMMs) struggle to handle hour-long video inputs due to the quadratic complexity of the causal self-attention operations, leading to high computational costs during training and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Weiming Ren , Wentao Ma , Huan Yang , Cong Wei , Ge Zhang , Wenhu Chen

Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention,…

Machine Learning · Computer Science 2024-06-03 Albert Gu , Tri Dao

U-shaped architectures have long dominated the field of medical image segmentation, while Transformers are widely employed for modeling long-range dependencies. The former typically handles scale variations implicitly by aggregating…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Yanhua Zhang , Ke Zhang , Jingyu Wang , Gabriella Balestra , Samanta Rosati , Yulin Wu , Wuwei Wang , Valentina Giannini

Mixture-of-Experts (MoE) has gained increasing popularity as a promising framework for scaling up large language models (LLMs). However, training MoE from scratch in a large-scale setting still suffers from data-hungry and instability…

Computation and Language · Computer Science 2024-06-25 Tong Zhu , Xiaoye Qu , Daize Dong , Jiacheng Ruan , Jingqi Tong , Conghui He , Yu Cheng

Transformer-based large language models (LLMs) are increasingly being adopted in networking research to address domain-specific challenges. However, their quadratic time complexity and substantial model sizes often result in significant…

Networking and Internet Architecture · Computer Science 2025-10-21 Linhan Xia , Mingzhan Yang , Jingjing Wang , Ziwei Yan , Yakun Ren , Guo Yu , Kai Lei

The deployment of large language models (LLMs) in real-world clinical applications is constrained by the fundamental trade-off between computational cost and the efficiency of linear-time models. To address this, we propose an LLM-based…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Hamad Khan , Saddam Hussain Khan

In recent years, Transformers have become the de-facto architecture for sequence modeling on text and a variety of multi-dimensional data, such as images and video. However, the use of self-attention layers in a Transformer incurs…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Shufan Li , Harkanwar Singh , Aditya Grover