English
Related papers

Related papers: Falcon Mamba: The First Competitive Attention-free…

200 papers

Large language models (LLMs) have made significant advances in complex reasoning tasks, yet they remain bottlenecked by two core challenges: architectural inefficiency due to reliance on Transformers, and a lack of structured fine-tuning…

Machine Learning · Computer Science 2025-05-29 Xueliang Zhao , Wei Wu , Lingpeng Kong

Linear RNN architectures, like Mamba, can be competitive with Transformer models in language modeling while having advantageous deployment characteristics. Given the focus on training large-scale Transformer models, we consider the…

Machine Learning · Computer Science 2025-06-30 Junxiong Wang , Daniele Paliotta , Avner May , Alexander M. Rush , Tri Dao

In recent years, the talking head generation has become a focal point for researchers. Considerable effort is being made to refine lip-sync motion, capture expressive facial expressions, generate natural head poses, and achieve high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Farzaneh Jafari , Stefano Berretti , Anup Basu

Foundation models learn transferable representations, motivating growing interest in their application to wireless systems. Existing wireless foundation models are predominantly based on transformer architectures, whose quadratic…

Signal Processing · Electrical Eng. & Systems 2026-03-30 Tomer Raviv , Nir Shlezinger

Multimodal large language models (MLLMs) have attracted widespread interest and have rich applications. However, the inherent attention mechanism in its Transformer structure requires quadratic complexity and results in expensive…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Yanyuan Qiao , Zheng Yu , Longteng Guo , Sihan Chen , Zijia Zhao , Mingzhen Sun , Qi Wu , Jing Liu

We introduce llama-embed-nemotron-8b, an open-weights text embedding model that achieves state-of-the-art performance on the Multilingual Massive Text Embedding Benchmark (MMTEB) leaderboard as of October 21, 2025. While recent models show…

Computation and Language · Computer Science 2025-11-11 Yauhen Babakhin , Radek Osmulski , Ronay Ak , Gabriel Moreira , Mengyao Xu , Benedikt Schifferer , Bo Liu , Even Oldridge

Recently developed large language models (LLMs) such as ChatGPT, Claude, and Llama have demonstrated impressive abilities, and even surpass human-level performance in several tasks. Despite their success, the resource-intensive demands of…

Computation and Language · Computer Science 2024-06-17 Jie Wu , Yufeng Zhu , Lei Shen , Xuqing Lu

Large language models (LLMs) have advanced significantly due to the attention mechanism, but their quadratic complexity and linear memory demands limit their performance on long-context tasks. Recently, researchers introduced Mamba, an…

Computation and Language · Computer Science 2024-10-22 Wangjie You , Zecheng Tang , Juntao Li , Lili Yao , Min Zhang

Mamba is an effective state space model with linear computation complexity. It has recently shown impressive efficiency in dealing with high-resolution inputs across various vision tasks. In this paper, we reveal that the powerful Mamba…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Dongchen Han , Ziyi Wang , Zhuofan Xia , Yizeng Han , Yifan Pu , Chunjiang Ge , Jun Song , Shiji Song , Bo Zheng , Gao Huang

Mamba, a special case of the State Space Model, is gaining popularity as an alternative to template-based deep learning approaches in medical image analysis. While transformers are powerful architectures, they have drawbacks, including…

Sequential recommendation systems aim to predict users' next preferences based on their interaction histories, but existing approaches face critical limitations in efficiency and multi-scale pattern recognition. While Transformer-based…

Information Retrieval · Computer Science 2025-05-08 Qianru Zhang , Liang Qu , Honggang Wen , Dong Huang , Siu-Ming Yiu , Nguyen Quoc Viet Hung , Hongzhi Yin

Transformers and Mamba, initially invented for natural language processing, have inspired backbone architectures for visual recognition. Recent studies integrated Local Attention Transformers with Mamba to capture both local details and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Meng Lou , Yunxiang Fu , Yizhou Yu

We introduce Nemotron-Nano-9B-v2, a hybrid Mamba-Transformer language model designed to increase throughput for reasoning workloads while achieving state-of-the-art accuracy compared to similarly-sized models. Nemotron-Nano-9B-v2 builds on…

Computation and Language · Computer Science 2025-09-03 NVIDIA , : , Aarti Basant , Abhijit Khairnar , Abhijit Paithankar , Abhinav Khattar , Adithya Renduchintala , Aditya Malte , Akhiad Bercovich , Akshay Hazare , Alejandra Rico , Aleksander Ficek , Alex Kondratenko , Alex Shaposhnikov , Alexander Bukharin , Ali Taghibakhshi , Amelia Barton , Ameya Sunil Mahabaleshwarkar , Amy Shen , Andrew Tao , Ann Guan , Anna Shors , Anubhav Mandarwal , Arham Mehta , Arun Venkatesan , Ashton Sharabiani , Ashwath Aithal , Ashwin Poojary , Ayush Dattagupta , Balaram Buddharaju , Banghua Zhu , Barnaby Simkin , Bilal Kartal , Bita Darvish Rouhani , Bobby Chen , Boris Ginsburg , Brandon Norick , Brian Yu , Bryan Catanzaro , Charles Wang , Charlie Truong , Chetan Mungekar , Chintan Patel , Chris Alexiuk , Christian Munley , Christopher Parisien , Dan Su , Daniel Afrimi , Daniel Korzekwa , Daniel Rohrer , Daria Gitman , David Mosallanezhad , Deepak Narayanan , Dima Rekesh , Dina Yared , Dmytro Pykhtar , Dong Ahn , Duncan Riach , Eileen Long , Elliott Ning , Eric Chung , Erick Galinkin , Evelina Bakhturina , Gargi Prasad , Gerald Shen , Haifeng Qian , Haim Elisha , Harsh Sharma , Hayley Ross , Helen Ngo , Herman Sahota , Hexin Wang , Hoo Chang Shin , Hua Huang , Iain Cunningham , Igor Gitman , Ivan Moshkov , Jaehun Jung , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jian Zhang , Jiaqi Zeng , Jimmy Zhang , Jinze Xue , Jocelyn Huang , Joey Conway , John Kamalu , Jonathan Cohen , Joseph Jennings , Julien Veron Vialard , Junkeun Yi , Jupinder Parmar , Kari Briski , Katherine Cheung , Katherine Luna , Keith Wyss , Keshav Santhanam , Kezhi Kong , Krzysztof Pawelec , Kumar Anik , Kunlun Li , Kushan Ahmadian , Lawrence McAfee , Laya Sleiman , Leon Derczynski , Luis Vega , Maer Rodrigues de Melo , Makesh Narsimhan Sreedhar , Marcin Chochowski , Mark Cai , Markus Kliegl , Marta Stepniewska-Dziubinska , Matvei Novikov , Mehrzad Samadi , Meredith Price , Meriem Boubdir , Michael Boone , Michael Evans , Michal Bien , Michal Zawalski , Miguel Martinez , Mike Chrzanowski , Mohammad Shoeybi , Mostofa Patwary , Namit Dhameja , Nave Assaf , Negar Habibi , Nidhi Bhatia , Nikki Pope , Nima Tajbakhsh , Nirmal Kumar Juluru , Oleg Rybakov , Oleksii Hrinchuk , Oleksii Kuchaiev , Oluwatobi Olabiyi , Pablo Ribalta , Padmavathy Subramanian , Parth Chadha , Pavlo Molchanov , Peter Dykas , Peter Jin , Piotr Bialecki , Piotr Januszewski , Pradeep Thalasta , Prashant Gaikwad , Prasoon Varshney , Pritam Gundecha , Przemek Tredak , Rabeeh Karimi Mahabadi , Rajen Patel , Ran El-Yaniv , Ranjit Rajan , Ria Cheruvu , Rima Shahbazyan , Ritika Borkar , Ritu Gala , Roger Waleffe , Ruoxi Zhang , Russell J. Hewett , Ryan Prenger , Sahil Jain , Samuel Kriman , Sanjeev Satheesh , Saori Kaji , Sarah Yurick , Saurav Muralidharan , Sean Narenthiran , Seonmyeong Bak , Sepehr Sameni , Seungju Han , Shanmugam Ramasamy , Shaona Ghosh , Sharath Turuvekere Sreenivas , Shelby Thomas , Shizhe Diao , Shreya Gopal , Shrimai Prabhumoye , Shubham Toshniwal , Shuoyang Ding , Siddharth Singh , Siddhartha Jain , Somshubra Majumdar , Soumye Singhal , Stefania Alborghetti , Syeda Nahida Akter , Terry Kong , Tim Moon , Tomasz Hliwiak , Tomer Asida , Tony Wang , Tugrul Konuk , Twinkle Vashishth , Tyler Poon , Udi Karpas , Vahid Noroozi , Venkat Srinivasan , Vijay Korthikanti , Vikram Fugro , Vineeth Kalluru , Vitaly Kurin , Vitaly Lavrukhin , Wasi Uddin Ahmad , Wei Du , Wonmin Byeon , Ximing Lu , Xin Dong , Yashaswi Karnati , Yejin Choi , Yian Zhang , Ying Lin , Yonggan Fu , Yoshi Suhara , Zhen Dong , Zhiyu Li , Zhongbo Zhu , Zijia Chen

Mamba extends earlier state space models (SSMs) by introducing input-dependent dynamics, and has demonstrated strong empirical performance across a range of domains, including language modeling, computer vision, and foundation models.…

Machine Learning · Computer Science 2025-05-15 Annan Yu , N. Benjamin Erichson

This study explores replacing Transformers in Visual Language Models (VLMs) with Mamba, a recent structured state space model (SSM) that demonstrates promising performance in sequence modeling. We test models up to 3B parameters under…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Georgios Pantazopoulos , Malvina Nikandrou , Alessandro Suglia , Oliver Lemon , Arash Eshghi

Multilingual automatic speech recognition (ASR) remains a challenging task, especially when balancing performance across high- and low-resource languages. Recent advances in sequence modeling suggest that architectures beyond Transformers…

Computation and Language · Computer Science 2025-10-24 Mohamed Nabih Ali , Daniele Falavigna , Alessio Brutti

This paper explores the capability of Mamba, a recently proposed architecture based on state space models (SSMs), as a competitive alternative to Transformer-based models. In the speech domain, well-designed Transformer-based models, such…

Sound · Computer Science 2024-06-25 Koichi Miyazaki , Yoshiki Masuyama , Masato Murata

Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-attention mechanism in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-29 Xiangyu Zhang , Qiquan Zhang , Hexin Liu , Tianyi Xiao , Xinyuan Qian , Beena Ahmed , Eliathamby Ambikairajah , Haizhou Li , Julien Epps

Transformers have revolutionized deep learning across various tasks, including audio representation learning, due to their powerful modeling capabilities. However, they often suffer from quadratic complexity in both GPU memory usage and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-06 Siavash Shams , Sukru Samet Dindar , Xilin Jiang , Nima Mesgarani

In this work, we introduce Gemma 2, a new addition to the Gemma family of lightweight, state-of-the-art open models, ranging in scale from 2 billion to 27 billion parameters. In this new version, we apply several known technical…

Computation and Language · Computer Science 2024-10-03 Gemma Team , Morgane Riviere , Shreya Pathak , Pier Giuseppe Sessa , Cassidy Hardin , Surya Bhupatiraju , Léonard Hussenot , Thomas Mesnard , Bobak Shahriari , Alexandre Ramé , Johan Ferret , Peter Liu , Pouya Tafti , Abe Friesen , Michelle Casbon , Sabela Ramos , Ravin Kumar , Charline Le Lan , Sammy Jerome , Anton Tsitsulin , Nino Vieillard , Piotr Stanczyk , Sertan Girgin , Nikola Momchev , Matt Hoffman , Shantanu Thakoor , Jean-Bastien Grill , Behnam Neyshabur , Olivier Bachem , Alanna Walton , Aliaksei Severyn , Alicia Parrish , Aliya Ahmad , Allen Hutchison , Alvin Abdagic , Amanda Carl , Amy Shen , Andy Brock , Andy Coenen , Anthony Laforge , Antonia Paterson , Ben Bastian , Bilal Piot , Bo Wu , Brandon Royal , Charlie Chen , Chintu Kumar , Chris Perry , Chris Welty , Christopher A. Choquette-Choo , Danila Sinopalnikov , David Weinberger , Dimple Vijaykumar , Dominika Rogozińska , Dustin Herbison , Elisa Bandy , Emma Wang , Eric Noland , Erica Moreira , Evan Senter , Evgenii Eltyshev , Francesco Visin , Gabriel Rasskin , Gary Wei , Glenn Cameron , Gus Martins , Hadi Hashemi , Hanna Klimczak-Plucińska , Harleen Batra , Harsh Dhand , Ivan Nardini , Jacinda Mein , Jack Zhou , James Svensson , Jeff Stanway , Jetha Chan , Jin Peng Zhou , Joana Carrasqueira , Joana Iljazi , Jocelyn Becker , Joe Fernandez , Joost van Amersfoort , Josh Gordon , Josh Lipschultz , Josh Newlan , Ju-yeong Ji , Kareem Mohamed , Kartikeya Badola , Kat Black , Katie Millican , Keelin McDonell , Kelvin Nguyen , Kiranbir Sodhia , Kish Greene , Lars Lowe Sjoesund , Lauren Usui , Laurent Sifre , Lena Heuermann , Leticia Lago , Lilly McNealus , Livio Baldini Soares , Logan Kilpatrick , Lucas Dixon , Luciano Martins , Machel Reid , Manvinder Singh , Mark Iverson , Martin Görner , Mat Velloso , Mateo Wirth , Matt Davidow , Matt Miller , Matthew Rahtz , Matthew Watson , Meg Risdal , Mehran Kazemi , Michael Moynihan , Ming Zhang , Minsuk Kahng , Minwoo Park , Mofi Rahman , Mohit Khatwani , Natalie Dao , Nenshad Bardoliwalla , Nesh Devanathan , Neta Dumai , Nilay Chauhan , Oscar Wahltinez , Pankil Botarda , Parker Barnes , Paul Barham , Paul Michel , Pengchong Jin , Petko Georgiev , Phil Culliton , Pradeep Kuppala , Ramona Comanescu , Ramona Merhej , Reena Jana , Reza Ardeshir Rokni , Rishabh Agarwal , Ryan Mullins , Samaneh Saadat , Sara Mc Carthy , Sarah Cogan , Sarah Perrin , Sébastien M. R. Arnold , Sebastian Krause , Shengyang Dai , Shruti Garg , Shruti Sheth , Sue Ronstrom , Susan Chan , Timothy Jordan , Ting Yu , Tom Eccles , Tom Hennigan , Tomas Kocisky , Tulsee Doshi , Vihan Jain , Vikas Yadav , Vilobh Meshram , Vishal Dharmadhikari , Warren Barkley , Wei Wei , Wenming Ye , Woohyun Han , Woosuk Kwon , Xiang Xu , Zhe Shen , Zhitao Gong , Zichuan Wei , Victor Cotruta , Phoebe Kirk , Anand Rao , Minh Giang , Ludovic Peran , Tris Warkentin , Eli Collins , Joelle Barral , Zoubin Ghahramani , Raia Hadsell , D. Sculley , Jeanine Banks , Anca Dragan , Slav Petrov , Oriol Vinyals , Jeff Dean , Demis Hassabis , Koray Kavukcuoglu , Clement Farabet , Elena Buchatskaya , Sebastian Borgeaud , Noah Fiedel , Armand Joulin , Kathleen Kenealy , Robert Dadashi , Alek Andreev