English
Related papers

Related papers: Jamba: A Hybrid Transformer-Mamba Language Model

200 papers

In recent years, Transformers-based models have made significant progress in the field of image restoration by leveraging their inherent ability to capture complex contextual features. Recently, Mamba models have made a splash in the field…

Computer Vision and Pattern Recognition · Computer Science 2025-02-10 Juan Wen , Weiyan Hou , Luc Van Gool , Radu Timofte

Large language models operate in distinct compute-bound prefill followed by memory bandwidth-bound decode phases. Hybrid Mamba-Transformer models inherit this asymmetry while adding state space model (SSM) recurrences and element-wise…

Hardware Architecture · Computer Science 2026-03-17 Alish Kanani , Sangwan Lee , Han Lyu , Jiahao Lin , Jaehyun Park , Umit Y. Ogras

We introduce MoMa, a novel modality-aware mixture-of-experts (MoE) architecture designed for pre-training mixed-modal, early-fusion language models. MoMa processes images and text in arbitrary sequences by dividing expert modules into…

Artificial Intelligence · Computer Science 2024-08-13 Xi Victoria Lin , Akshat Shrivastava , Liang Luo , Srinivasan Iyer , Mike Lewis , Gargi Ghosh , Luke Zettlemoyer , Armen Aghajanyan

Self attention encoders such as Bidirectional Encoder Representations from Transformers(BERT) scale quadratically with sequence length, making long context modeling expensive. Linear time state space models, such as Mamba, are efficient;…

Computation and Language · Computer Science 2026-03-04 Jinwoong Kim , Sangjin Park

With the explosive growth of data, long-sequence modeling has become increasingly important in tasks such as natural language processing and bioinformatics. However, existing methods face inherent trade-offs between efficiency and memory.…

Machine Learning · Computer Science 2025-10-07 Youjin Wang , Yangjingyi Chen , Jiahao Yan , Jiaxuan Lu , Xiao Sun

Understanding videos is one of the fundamental directions in computer vision research, with extensive efforts dedicated to exploring various architectures such as RNN, 3D CNN, and Transformers. The newly proposed architecture of state space…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Guo Chen , Yifei Huang , Jilan Xu , Baoqi Pei , Zhe Chen , Zhiqi Li , Jiahao Wang , Kunchang Li , Tong Lu , Limin Wang

Mamba extends earlier state space models (SSMs) by introducing input-dependent dynamics, and has demonstrated strong empirical performance across a range of domains, including language modeling, computer vision, and foundation models.…

Machine Learning · Computer Science 2025-05-15 Annan Yu , N. Benjamin Erichson

In scene text detection, Transformer-based methods have addressed the global feature extraction limitations inherent in traditional convolution neural network-based methods. However, most directly rely on native Transformer attention layers…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Qiyan Zhao , Yue Yan , Da-Han Wang

Cracks pose safety risks to infrastructure and cannot be overlooked. The prevailing structures in existing crack segmentation networks predominantly consist of CNNs or Transformers. However, CNNs exhibit a deficiency in global modeling…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Zhili He , Yu-Hsing Wang

Fine-tuning Large Language Models (LLMs) is a common practice to adapt pre-trained models for specific applications. While methods like LoRA have effectively addressed GPU memory constraints during fine-tuning, their performance often falls…

Computation and Language · Computer Science 2024-07-23 Dengchun Li , Yingzi Ma , Naizheng Wang , Zhengmao Ye , Zhiyuan Cheng , Yinghao Tang , Yan Zhang , Lei Duan , Jie Zuo , Cal Yang , Mingjie Tang

The Mixture of Experts (MoE) architecture is a cornerstone of modern state-of-the-art (SOTA) large language models (LLMs). MoE models facilitate scalability by enabling sparse parameter activation. However, traditional MoE architecture uses…

Computation and Language · Computer Science 2025-08-12 Haoyuan Wu , Haoxing Chen , Xiaodong Chen , Zhanchao Zhou , Tieyuan Chen , Yihong Zhuang , Guoshan Lu , Zenan Huang , Junbo Zhao , Lin Liu , Zhenzhong Lan , Bei Yu , Jianguo Li

We introduce Nemotron-Nano-9B-v2, a hybrid Mamba-Transformer language model designed to increase throughput for reasoning workloads while achieving state-of-the-art accuracy compared to similarly-sized models. Nemotron-Nano-9B-v2 builds on…

Computation and Language · Computer Science 2025-09-03 NVIDIA , : , Aarti Basant , Abhijit Khairnar , Abhijit Paithankar , Abhinav Khattar , Adithya Renduchintala , Aditya Malte , Akhiad Bercovich , Akshay Hazare , Alejandra Rico , Aleksander Ficek , Alex Kondratenko , Alex Shaposhnikov , Alexander Bukharin , Ali Taghibakhshi , Amelia Barton , Ameya Sunil Mahabaleshwarkar , Amy Shen , Andrew Tao , Ann Guan , Anna Shors , Anubhav Mandarwal , Arham Mehta , Arun Venkatesan , Ashton Sharabiani , Ashwath Aithal , Ashwin Poojary , Ayush Dattagupta , Balaram Buddharaju , Banghua Zhu , Barnaby Simkin , Bilal Kartal , Bita Darvish Rouhani , Bobby Chen , Boris Ginsburg , Brandon Norick , Brian Yu , Bryan Catanzaro , Charles Wang , Charlie Truong , Chetan Mungekar , Chintan Patel , Chris Alexiuk , Christian Munley , Christopher Parisien , Dan Su , Daniel Afrimi , Daniel Korzekwa , Daniel Rohrer , Daria Gitman , David Mosallanezhad , Deepak Narayanan , Dima Rekesh , Dina Yared , Dmytro Pykhtar , Dong Ahn , Duncan Riach , Eileen Long , Elliott Ning , Eric Chung , Erick Galinkin , Evelina Bakhturina , Gargi Prasad , Gerald Shen , Haifeng Qian , Haim Elisha , Harsh Sharma , Hayley Ross , Helen Ngo , Herman Sahota , Hexin Wang , Hoo Chang Shin , Hua Huang , Iain Cunningham , Igor Gitman , Ivan Moshkov , Jaehun Jung , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jian Zhang , Jiaqi Zeng , Jimmy Zhang , Jinze Xue , Jocelyn Huang , Joey Conway , John Kamalu , Jonathan Cohen , Joseph Jennings , Julien Veron Vialard , Junkeun Yi , Jupinder Parmar , Kari Briski , Katherine Cheung , Katherine Luna , Keith Wyss , Keshav Santhanam , Kezhi Kong , Krzysztof Pawelec , Kumar Anik , Kunlun Li , Kushan Ahmadian , Lawrence McAfee , Laya Sleiman , Leon Derczynski , Luis Vega , Maer Rodrigues de Melo , Makesh Narsimhan Sreedhar , Marcin Chochowski , Mark Cai , Markus Kliegl , Marta Stepniewska-Dziubinska , Matvei Novikov , Mehrzad Samadi , Meredith Price , Meriem Boubdir , Michael Boone , Michael Evans , Michal Bien , Michal Zawalski , Miguel Martinez , Mike Chrzanowski , Mohammad Shoeybi , Mostofa Patwary , Namit Dhameja , Nave Assaf , Negar Habibi , Nidhi Bhatia , Nikki Pope , Nima Tajbakhsh , Nirmal Kumar Juluru , Oleg Rybakov , Oleksii Hrinchuk , Oleksii Kuchaiev , Oluwatobi Olabiyi , Pablo Ribalta , Padmavathy Subramanian , Parth Chadha , Pavlo Molchanov , Peter Dykas , Peter Jin , Piotr Bialecki , Piotr Januszewski , Pradeep Thalasta , Prashant Gaikwad , Prasoon Varshney , Pritam Gundecha , Przemek Tredak , Rabeeh Karimi Mahabadi , Rajen Patel , Ran El-Yaniv , Ranjit Rajan , Ria Cheruvu , Rima Shahbazyan , Ritika Borkar , Ritu Gala , Roger Waleffe , Ruoxi Zhang , Russell J. Hewett , Ryan Prenger , Sahil Jain , Samuel Kriman , Sanjeev Satheesh , Saori Kaji , Sarah Yurick , Saurav Muralidharan , Sean Narenthiran , Seonmyeong Bak , Sepehr Sameni , Seungju Han , Shanmugam Ramasamy , Shaona Ghosh , Sharath Turuvekere Sreenivas , Shelby Thomas , Shizhe Diao , Shreya Gopal , Shrimai Prabhumoye , Shubham Toshniwal , Shuoyang Ding , Siddharth Singh , Siddhartha Jain , Somshubra Majumdar , Soumye Singhal , Stefania Alborghetti , Syeda Nahida Akter , Terry Kong , Tim Moon , Tomasz Hliwiak , Tomer Asida , Tony Wang , Tugrul Konuk , Twinkle Vashishth , Tyler Poon , Udi Karpas , Vahid Noroozi , Venkat Srinivasan , Vijay Korthikanti , Vikram Fugro , Vineeth Kalluru , Vitaly Kurin , Vitaly Lavrukhin , Wasi Uddin Ahmad , Wei Du , Wonmin Byeon , Ximing Lu , Xin Dong , Yashaswi Karnati , Yejin Choi , Yian Zhang , Ying Lin , Yonggan Fu , Yoshi Suhara , Zhen Dong , Zhiyu Li , Zhongbo Zhu , Zijia Chen

The Sparsely-Activated Mixture-of-Experts (MoE) has gained increasing popularity for scaling up large language models (LLMs) without exploding computational costs. Despite its success, the current design faces a challenge where all experts…

Machine Learning · Computer Science 2024-09-20 Manxi Sun , Wei Liu , Jian Luan , Pengzhi Gao , Bin Wang

Modern multivariate time series forecasting primarily relies on two architectures: the Transformer with attention mechanism and Mamba. In natural language processing, an approach has been used that combines local window attention for…

Machine Learning · Computer Science 2025-09-26 Itay Katav , Aryeh Kontorovich

State Space Models (SSMs), particularly the Mamba architecture, have recently emerged as powerful alternatives to Transformers for sequence modeling, offering linear computational complexity while achieving competitive performance. Yet,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Mohamed A. Mabrok , Yalda Zafari

Mixture-of-Experts (MoE) has emerged as a practical approach to scale up parameters for the Transformer model to achieve better generalization while maintaining a sub-linear increase in computation overhead. Current MoE models are mainly…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-03 Shuqing Luo , Jie Peng , Pingzhi Li , Hanrui Wang , Tianlong Chen

Recent advances in language modeling have demonstrated the effectiveness of State Space Models (SSMs) for efficient sequence modeling. While hybrid architectures such as Samba and the decoder-decoder architecture, YOCO, have shown promising…

In the past decade, Convolutional Neural Networks (CNNs) and Transformers have achieved wide applicaiton in semantic segmentation tasks. Although CNNs with Transformer models greatly improve performance, the global context modeling remains…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Feixiang Du , Shengkun Wu

Large language models (LLMs) are widely used but expensive to run, especially as inference workloads grow. To lower costs, maximizing the request batch size by managing GPU memory efficiently is crucial. While PagedAttention has recently…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-03-25 Chen Zhang , Kuntai Du , Shu Liu , Woosuk Kwon , Xiangxi Mo , Yufeng Wang , Xiaoxuan Liu , Kaichao You , Zhuohan Li , Mingsheng Long , Jidong Zhai , Joseph Gonzalez , Ion Stoica

Within the family of convolutional neural networks, InceptionNeXt has shown excellent competitiveness in image classification and a number of downstream tasks. Built on parallel one-dimensional strip convolutions, however, it suffers from…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Yuhang Wang , Jun Li , Zhijian Wu , Jifeng Shen , Jianhua Xu , Wankou Yang