English
Related papers

Related papers: Falcon Mamba: The First Competitive Attention-free…

200 papers

Image generation models have encountered challenges related to scalability and quadratic complexity, primarily due to the reliance on Transformer-based backbones. In this study, we introduce MaskMamba, a novel hybrid model that combines…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Wenchao Chen , Liqiang Niu , Ziyao Lu , Fandong Meng , Jie Zhou

The Mamba model has gained significant attention for its computational advantages over Transformer-based models, while achieving comparable performance across a wide range of language tasks. Like Transformers, Mamba exhibits in-context…

Machine Learning · Computer Science 2025-10-02 Hongkang Li , Songtao Lu , Xiaodong Cui , Pin-Yu Chen , Meng Wang

In recent years, the application of multimodal large language models (MLLM) in various fields has achieved remarkable success. However, as the foundation model for many downstream tasks, current MLLMs are composed of the well-known…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Han Zhao , Min Zhang , Wei Zhao , Pengxiang Ding , Siteng Huang , Donglin Wang

As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemotron-H, a family of 8B and 56B/47B hybrid Mamba-Transformer…

Computation and Language · Computer Science 2025-09-09 NVIDIA , : , Aaron Blakeman , Aarti Basant , Abhinav Khattar , Adithya Renduchintala , Akhiad Bercovich , Aleksander Ficek , Alexis Bjorlin , Ali Taghibakhshi , Amala Sanjay Deshmukh , Ameya Sunil Mahabaleshwarkar , Andrew Tao , Anna Shors , Ashwath Aithal , Ashwin Poojary , Ayush Dattagupta , Balaram Buddharaju , Bobby Chen , Boris Ginsburg , Boxin Wang , Brandon Norick , Brian Butterfield , Bryan Catanzaro , Carlo del Mundo , Chengyu Dong , Christine Harvey , Christopher Parisien , Dan Su , Daniel Korzekwa , Danny Yin , Daria Gitman , David Mosallanezhad , Deepak Narayanan , Denys Fridman , Dima Rekesh , Ding Ma , Dmytro Pykhtar , Dong Ahn , Duncan Riach , Dusan Stosic , Eileen Long , Elad Segal , Ellie Evans , Eric Chung , Erick Galinkin , Evelina Bakhturina , Ewa Dobrowolska , Fei Jia , Fuxiao Liu , Gargi Prasad , Gerald Shen , Guilin Liu , Guo Chen , Haifeng Qian , Helen Ngo , Hongbin Liu , Hui Li , Igor Gitman , Ilia Karmanov , Ivan Moshkov , Izik Golan , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jarno Seppanen , Jason Lu , Jason Sewall , Jiaqi Zeng , Jiaxuan You , Jimmy Zhang , Jing Zhang , Jining Huang , Jinze Xue , Jocelyn Huang , Joey Conway , John Kamalu , Jon Barker , Jonathan Cohen , Joseph Jennings , Jupinder Parmar , Karan Sapra , Kari Briski , Kateryna Chumachenko , Katherine Luna , Keshav Santhanam , Kezhi Kong , Kirthi Sivamani , Krzysztof Pawelec , Kumar Anik , Kunlun Li , Lawrence McAfee , Leon Derczynski , Lindsey Pavao , Luis Vega , Lukas Voegtle , Maciej Bala , Maer Rodrigues de Melo , Makesh Narsimhan Sreedhar , Marcin Chochowski , Markus Kliegl , Marta Stepniewska-Dziubinska , Matthieu Le , Matvei Novikov , Mehrzad Samadi , Michael Andersch , Michael Evans , Miguel Martinez , Mike Chrzanowski , Mike Ranzinger , Mikolaj Blaz , Misha Smelyanskiy , Mohamed Fawzy , Mohammad Shoeybi , Mostofa Patwary , Nayeon Lee , Nima Tajbakhsh , Ning Xu , Oleg Rybakov , Oleksii Kuchaiev , Olivier Delalleau , Osvald Nitski , Parth Chadha , Pasha Shamis , Paulius Micikevicius , Pavlo Molchanov , Peter Dykas , Philipp Fischer , Pierre-Yves Aquilanti , Piotr Bialecki , Prasoon Varshney , Pritam Gundecha , Przemek Tredak , Rabeeh Karimi , Rahul Kandu , Ran El-Yaniv , Raviraj Joshi , Roger Waleffe , Ruoxi Zhang , Sabrina Kavanaugh , Sahil Jain , Samuel Kriman , Sangkug Lym , Sanjeev Satheesh , Saurav Muralidharan , Sean Narenthiran , Selvaraj Anandaraj , Seonmyeong Bak , Sergey Kashirsky , Seungju Han , Shantanu Acharya , Shaona Ghosh , Sharath Turuvekere Sreenivas , Sharon Clay , Shelby Thomas , Shrimai Prabhumoye , Shubham Pachori , Shubham Toshniwal , Shyamala Prayaga , Siddhartha Jain , Sirshak Das , Slawek Kierat , Somshubra Majumdar , Song Han , Soumye Singhal , Sriharsha Niverty , Stefania Alborghetti , Suseella Panguluri , Swetha Bhendigeri , Syeda Nahida Akter , Szymon Migacz , Tal Shiri , Terry Kong , Timo Roman , Tomer Ronen , Trisha Saar , Tugrul Konuk , Tuomas Rintamaki , Tyler Poon , Ushnish De , Vahid Noroozi , Varun Singh , Vijay Korthikanti , Vitaly Kurin , Wasi Uddin Ahmad , Wei Du , Wei Ping , Wenliang Dai , Wonmin Byeon , Xiaowei Ren , Yao Xu , Yejin Choi , Yian Zhang , Ying Lin , Yoshi Suhara , Zhiding Yu , Zhiqi Li , Zhiyu Li , Zhongbo Zhu , Zhuolin Yang , Zijia Chen

With the growing demand for deploying large language models (LLMs) across diverse applications, improving their inference efficiency is crucial for sustainable and democratized access. However, retraining LLMs to meet new user-specific…

Machine Learning · Computer Science 2026-01-21 Mingyu Yang , Mehdi Rezagholizadeh , Guihong Li , Vikram Appia , Emad Barsoum

The Transformer architecture has shown a remarkable ability in modeling global relationships. However, it poses a significant computational challenge when processing high-dimensional medical images. This hinders its development and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Zhaohu Xing , Tian Ye , Yijun Yang , Guang Liu , Lei Zhu

Previous research on lightweight models has primarily focused on CNNs and Transformer-based designs. CNNs, with their local receptive fields, struggle to capture long-range dependencies, while Transformers, despite their global modeling…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Haoyang He , Jiangning Zhang , Yuxuan Cai , Hongxu Chen , Xiaobin Hu , Zhenye Gan , Yabiao Wang , Chengjie Wang , Yunsheng Wu , Lei Xie

Transformer, a deep neural network architecture, has long dominated the field of natural language processing and beyond. Nevertheless, the recent introduction of Mamba challenges its supremacy, sparks considerable interest among…

Computation and Language · Computer Science 2024-06-25 Yuchen Zou , Yineng Chen , Zuchao Li , Lefei Zhang , Hai Zhao

We propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates transformer attention mechanisms with state space models (SSMs) for enhanced efficiency. Attention heads provide…

Transformer-based models have become increasingly popular and have impacted speech-processing research owing to their exceptional performance in sequence modeling. Recently, a promising model architecture, Mamba, has emerged as a potential…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-27 Wen-Yuan Ting , Wenze Ren , Rong Chao , Hsin-Yi Lin , Yu Tsao , Fan-Gang Zeng

We introduce Motif-2-12.7B, a new open-weight foundation model that pushes the efficiency frontier of large language models by combining architectural innovation with system-level optimization. Designed for scalable language understanding…

Large language models (LLMs) have transformed NLP, yet their integration with audio remains underexplored despite audio's centrality to human communication. We introduce Falcon3-Audio, a family of Audio-Language Models (ALMs) built on…

Efficiently modeling sequences with infinite context length has long been a challenging problem. Previous approaches have either suffered from quadratic computational complexity or limited extrapolation ability in length generalization. In…

Computation and Language · Computer Science 2025-03-03 Liliang Ren , Yang Liu , Yadong Lu , Yelong Shen , Chen Liang , Weizhu Chen

Time series foundation models have demonstrated strong performance in zero-shot learning, making them well-suited for predicting rapidly evolving patterns in real-world applications where relevant training data are scarce. However, most of…

Machine Learning · Computer Science 2024-11-06 Haoyu Ma , Yushu Chen , Wenlai Zhao , Jinzhe Yang , Yingsheng Ji , Xinghua Xu , Xiaozhu Liu , Hao Jing , Shengzhuo Liu , Guangwen Yang

In recent years, robust matching methods using deep learning-based approaches have been actively studied and improved in computer vision tasks. However, there remains a persistent demand for both robust and fast matching techniques. To…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Kihwan Ryoo , Hyungtae Lim , Hyun Myung

Recent work has shown that attention-based language models excel at recall, the ability to ground generations in tokens previously seen in context. However, the efficiency of attention-based models is bottle-necked during inference by the…

Computation and Language · Computer Science 2025-03-10 Simran Arora , Sabri Eyuboglu , Michael Zhang , Aman Timalsina , Silas Alberti , Dylan Zinsley , James Zou , Atri Rudra , Christopher Ré

Transformers are the current architecture of choice for NLP, but their attention layers do not scale well to long contexts. Recent works propose to replace attention with linear recurrent layers -- this is the case for state space models,…

Computation and Language · Computer Science 2024-07-09 Hugo Pitorro , Pavlo Vasylenko , Marcos Treviso , André F. T. Martins

We train a suite of multimodal foundation models (MMFM) using the popular LLaVA framework with the recently released Gemma family of large language models (LLMs). Of particular interest is the 2B parameter Gemma model, which provides…

Computation and Language · Computer Science 2024-06-12 Musashi Hinck , Matthew L. Olson , David Cobbley , Shao-Yen Tseng , Vasudev Lal

We introduce the Bittensor Language Model, called "BTLM-3B-8K", a new state-of-the-art 3 billion parameter open-source language model. BTLM-3B-8K was trained on 627B tokens from the SlimPajama dataset with a mixture of 2,048 and 8,192…

The Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech processing, the Mamba-based model's performance varies…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-12 Xiangyu Zhang , Jianbo Ma , Mostafa Shahin , Beena Ahmed , Julien Epps