English
Related papers

Related papers: Jamba-1.5: Hybrid Transformer-Mamba Models at Scal…

200 papers

Multimodal Large Language Models (MLLMs) have attracted much attention for their multifunctionality. However, traditional Transformer architectures incur significant overhead due to their secondary computational complexity. To address this…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Wenjun Huang , Jiakai Pan , Jiahao Tang , Yanyu Ding , Yifei Xing , Yuhe Wang , Zhengzhuo Wang , Jianguo Hu

The deployment of large language models (LLMs) in real-world clinical applications is constrained by the fundamental trade-off between computational cost and the efficiency of linear-time models. To address this, we propose an LLM-based…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Hamad Khan , Saddam Hussain Khan

Transformer-based large language models (LLMs) are increasingly being adopted in networking research to address domain-specific challenges. However, their quadratic time complexity and substantial model sizes often result in significant…

Networking and Internet Architecture · Computer Science 2025-10-21 Linhan Xia , Mingzhan Yang , Jingjing Wang , Ziwei Yan , Yakun Ren , Guo Yu , Kai Lei

Transformers are the current architecture of choice for NLP, but their attention layers do not scale well to long contexts. Recent works propose to replace attention with linear recurrent layers -- this is the case for state space models,…

Computation and Language · Computer Science 2024-07-09 Hugo Pitorro , Pavlo Vasylenko , Marcos Treviso , André F. T. Martins

State space models (SSMs) have emerged as an efficient alternative to Transformer models for language modeling, offering linear computational complexity and constant memory usage as context length increases. However, despite their…

Computation and Language · Computer Science 2025-04-23 Zhifan Ye , Kejing Xia , Yonggan Fu , Xin Dong , Jihoon Hong , Xiangchi Yuan , Shizhe Diao , Jan Kautz , Pavlo Molchanov , Yingyan Celine Lin

In recent years, Transformers have become the de-facto architecture for sequence modeling on text and a variety of multi-dimensional data, such as images and video. However, the use of self-attention layers in a Transformer incurs…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Shufan Li , Harkanwar Singh , Aditya Grover

The diffusion model has long been plagued by scalability and quadratic complexity issues, especially within transformer-based structures. In this study, we aim to leverage the long sequence modeling capability of a State-Space Model called…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Vincent Tao Hu , Stefan Andreas Baumann , Ming Gui , Olga Grebenkova , Pingchuan Ma , Johannes Schusterbauer , Björn Ommer

Long-range sequence processing poses a significant challenge for Transformers due to their quadratic complexity in input length. A promising alternative is Mamba, which demonstrates high performance and achieves Transformer-level…

Machine Learning · Computer Science 2025-04-11 Assaf Ben-Kish , Itamar Zimerman , Shady Abu-Hussein , Nadav Cohen , Amir Globerson , Lior Wolf , Raja Giryes

Recent Mamba-based architectures for video understanding demonstrate promising computational efficiency and competitive performance, yet struggle with overfitting issues that hinder their scalability. To overcome this challenge, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yunze Liu , Peiran Wu , Cheng Liang , Junxiao Shen , Limin Wang , Li Yi

Large language models (LLMs) have advanced significantly due to the attention mechanism, but their quadratic complexity and linear memory demands limit their performance on long-context tasks. Recently, researchers introduced Mamba, an…

Computation and Language · Computer Science 2024-10-22 Wangjie You , Zecheng Tang , Juntao Li , Lili Yao , Min Zhang

Transformer-based models have become increasingly popular and have impacted speech-processing research owing to their exceptional performance in sequence modeling. Recently, a promising model architecture, Mamba, has emerged as a potential…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-27 Wen-Yuan Ting , Wenze Ren , Rong Chao , Hsin-Yi Lin , Yu Tsao , Fan-Gang Zeng

U-shaped architectures have long dominated the field of medical image segmentation, while Transformers are widely employed for modeling long-range dependencies. The former typically handles scale variations implicitly by aggregating…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Yanhua Zhang , Ke Zhang , Jingyu Wang , Gabriella Balestra , Samanta Rosati , Yulin Wu , Wuwei Wang , Valentina Giannini

Transformers and Mamba, initially invented for natural language processing, have inspired backbone architectures for visual recognition. Recent studies integrated Local Attention Transformers with Mamba to capture both local details and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Meng Lou , Yunxiang Fu , Yizhou Yu

Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead. We…

Machine Learning · Computer Science 2026-03-02 Vaibhav Singh , Oleksiy Ostapenko , Pierre-André Noël , Eugene Belilovsky , Torsten Scholak

In this work, we introduce ChatQA 2, an Llama 3.0-based model with a 128K context window, designed to bridge the gap between open-source LLMs and leading proprietary models (e.g., GPT-4-Turbo-2024-04-09) in long context understanding and…

Computation and Language · Computer Science 2025-02-18 Peng Xu , Wei Ping , Xianchao Wu , Chejian Xu , Zihan Liu , Mohammad Shoeybi , Bryan Catanzaro

As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemotron-H, a family of 8B and 56B/47B hybrid Mamba-Transformer…

Computation and Language · Computer Science 2025-09-09 NVIDIA , : , Aaron Blakeman , Aarti Basant , Abhinav Khattar , Adithya Renduchintala , Akhiad Bercovich , Aleksander Ficek , Alexis Bjorlin , Ali Taghibakhshi , Amala Sanjay Deshmukh , Ameya Sunil Mahabaleshwarkar , Andrew Tao , Anna Shors , Ashwath Aithal , Ashwin Poojary , Ayush Dattagupta , Balaram Buddharaju , Bobby Chen , Boris Ginsburg , Boxin Wang , Brandon Norick , Brian Butterfield , Bryan Catanzaro , Carlo del Mundo , Chengyu Dong , Christine Harvey , Christopher Parisien , Dan Su , Daniel Korzekwa , Danny Yin , Daria Gitman , David Mosallanezhad , Deepak Narayanan , Denys Fridman , Dima Rekesh , Ding Ma , Dmytro Pykhtar , Dong Ahn , Duncan Riach , Dusan Stosic , Eileen Long , Elad Segal , Ellie Evans , Eric Chung , Erick Galinkin , Evelina Bakhturina , Ewa Dobrowolska , Fei Jia , Fuxiao Liu , Gargi Prasad , Gerald Shen , Guilin Liu , Guo Chen , Haifeng Qian , Helen Ngo , Hongbin Liu , Hui Li , Igor Gitman , Ilia Karmanov , Ivan Moshkov , Izik Golan , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jarno Seppanen , Jason Lu , Jason Sewall , Jiaqi Zeng , Jiaxuan You , Jimmy Zhang , Jing Zhang , Jining Huang , Jinze Xue , Jocelyn Huang , Joey Conway , John Kamalu , Jon Barker , Jonathan Cohen , Joseph Jennings , Jupinder Parmar , Karan Sapra , Kari Briski , Kateryna Chumachenko , Katherine Luna , Keshav Santhanam , Kezhi Kong , Kirthi Sivamani , Krzysztof Pawelec , Kumar Anik , Kunlun Li , Lawrence McAfee , Leon Derczynski , Lindsey Pavao , Luis Vega , Lukas Voegtle , Maciej Bala , Maer Rodrigues de Melo , Makesh Narsimhan Sreedhar , Marcin Chochowski , Markus Kliegl , Marta Stepniewska-Dziubinska , Matthieu Le , Matvei Novikov , Mehrzad Samadi , Michael Andersch , Michael Evans , Miguel Martinez , Mike Chrzanowski , Mike Ranzinger , Mikolaj Blaz , Misha Smelyanskiy , Mohamed Fawzy , Mohammad Shoeybi , Mostofa Patwary , Nayeon Lee , Nima Tajbakhsh , Ning Xu , Oleg Rybakov , Oleksii Kuchaiev , Olivier Delalleau , Osvald Nitski , Parth Chadha , Pasha Shamis , Paulius Micikevicius , Pavlo Molchanov , Peter Dykas , Philipp Fischer , Pierre-Yves Aquilanti , Piotr Bialecki , Prasoon Varshney , Pritam Gundecha , Przemek Tredak , Rabeeh Karimi , Rahul Kandu , Ran El-Yaniv , Raviraj Joshi , Roger Waleffe , Ruoxi Zhang , Sabrina Kavanaugh , Sahil Jain , Samuel Kriman , Sangkug Lym , Sanjeev Satheesh , Saurav Muralidharan , Sean Narenthiran , Selvaraj Anandaraj , Seonmyeong Bak , Sergey Kashirsky , Seungju Han , Shantanu Acharya , Shaona Ghosh , Sharath Turuvekere Sreenivas , Sharon Clay , Shelby Thomas , Shrimai Prabhumoye , Shubham Pachori , Shubham Toshniwal , Shyamala Prayaga , Siddhartha Jain , Sirshak Das , Slawek Kierat , Somshubra Majumdar , Song Han , Soumye Singhal , Sriharsha Niverty , Stefania Alborghetti , Suseella Panguluri , Swetha Bhendigeri , Syeda Nahida Akter , Szymon Migacz , Tal Shiri , Terry Kong , Timo Roman , Tomer Ronen , Trisha Saar , Tugrul Konuk , Tuomas Rintamaki , Tyler Poon , Ushnish De , Vahid Noroozi , Varun Singh , Vijay Korthikanti , Vitaly Kurin , Wasi Uddin Ahmad , Wei Du , Wei Ping , Wenliang Dai , Wonmin Byeon , Xiaowei Ren , Yao Xu , Yejin Choi , Yian Zhang , Ying Lin , Yoshi Suhara , Zhiding Yu , Zhiqi Li , Zhiyu Li , Zhongbo Zhu , Zhuolin Yang , Zijia Chen

Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models such as Mamba has performed well on long-sequence modeling…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-30 Xiaoxue Gao , Nancy F. Chen

Recent works have shown the remarkable superiority of transformer models in reinforcement learning (RL), where the decision-making problem is formulated as sequential generation. Transformer-based agents could emerge with self-improvement…

Machine Learning · Computer Science 2024-06-04 Sili Huang , Jifeng Hu , Zhejian Yang , Liwei Yang , Tao Luo , Hechang Chen , Lichao Sun , Bo Yang

Extractive summarization of long documents is bottlenecked by quadratic complexity, often forcing truncation and limiting deployment in resource-constrained settings. We introduce the first Mamba-Transformer hybrid for extractive…

Computation and Language · Computer Science 2026-03-03 Nisrine Ait Khayi

Handling lengthy context is crucial for enhancing the recognition and understanding capabilities of multimodal large language models (MLLMs) in applications such as processing high-resolution images or high frame rate videos. The rise in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Jianing Zhou , Han Li , Shuai Zhang , Ning Xie , Ruijie Wang , Xiaohan Nie , Sheng Liu , Lingyun Wang