English
Related papers

Related papers: The Zamba2 Suite: Technical Report

200 papers

In this report, we introduce Falcon-H1, a new series of large language models (LLMs) featuring hybrid architecture designs optimized for both high performance and efficiency across diverse use cases. Unlike earlier Falcon models built…

Multi-modality image fusion aims to integrate the merits of images from different sources and render high-quality fusion images. However, existing feature extraction and fusion methods are either constrained by inherent local reduction bias…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Chenguang Zhu , Shan Gao , Huafeng Chen , Guangqian Guo , Chaowei Wang , Yaoxing Wang , Chen Shu Lei , Quanjiang Fan

Transformers are the current architecture of choice for NLP, but their attention layers do not scale well to long contexts. Recent works propose to replace attention with linear recurrent layers -- this is the case for state space models,…

Computation and Language · Computer Science 2024-07-09 Hugo Pitorro , Pavlo Vasylenko , Marcos Treviso , André F. T. Martins

Transformers bring significantly improved performance to the light field image super-resolution task due to their long-range dependency modeling capability. However, the inherently high computational complexity of their core self-attention…

Image and Video Processing · Electrical Eng. & Systems 2025-03-26 Zeqiang Wei , Kai Jin , Zeyi Hou , Kuan Song , Xiuzhuang Zhou

Efficiently modeling sequences with infinite context length has long been a challenging problem. Previous approaches have either suffered from quadratic computational complexity or limited extrapolation ability in length generalization. In…

Computation and Language · Computer Science 2025-03-03 Liliang Ren , Yang Liu , Yadong Lu , Yelong Shen , Chen Liang , Weizhu Chen

Existing state-of-the-art feature matchers capture long-range dependencies with Transformers but are hindered by high spatial complexity, leading to demanding training and highlatency inference. Striking a better balance between performance…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Xiaoyong Lu , Songlin Du

We introduce Mistral 7B v0.1, a 7-billion-parameter language model engineered for superior performance and efficiency. Mistral 7B outperforms Llama 2 13B across all evaluated benchmarks, and Llama 1 34B in reasoning, mathematics, and code…

We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by…

Computation and Language · Computer Science 2025-12-25 NVIDIA , : , Aaron Blakeman , Aaron Grattafiori , Aarti Basant , Abhibha Gupta , Abhinav Khattar , Adi Renduchintala , Aditya Vavre , Akanksha Shukla , Akhiad Bercovich , Aleksander Ficek , Aleksandr Shaposhnikov , Alex Kondratenko , Alexander Bukharin , Alexandre Milesi , Ali Taghibakhshi , Alisa Liu , Amelia Barton , Ameya Sunil Mahabaleshwarkar , Amir Klein , Amit Zuker , Amnon Geifman , Amy Shen , Anahita Bhiwandiwalla , Andrew Tao , Ann Guan , Anubhav Mandarwal , Arham Mehta , Ashwath Aithal , Ashwin Poojary , Asif Ahamed , Asma Kuriparambil Thekkumpate , Ayush Dattagupta , Banghua Zhu , Bardiya Sadeghi , Barnaby Simkin , Ben Lanir , Benedikt Schifferer , Besmira Nushi , Bilal Kartal , Bita Darvish Rouhani , Boris Ginsburg , Brandon Norick , Brandon Soubasis , Branislav Kisacanin , Brian Yu , Bryan Catanzaro , Carlo del Mundo , Chantal Hwang , Charles Wang , Cheng-Ping Hsieh , Chenghao Zhang , Chenhan Yu , Chetan Mungekar , Chintan Patel , Chris Alexiuk , Christopher Parisien , Collin Neale , Damon Mosk-Aoyama , Dan Su , Dane Corneil , Daniel Afrimi , Daniel Rohrer , Daniel Serebrenik , Daria Gitman , Daria Levy , Darko Stosic , David Mosallanezhad , Deepak Narayanan , Dhruv Nathawani , Dima Rekesh , Dina Yared , Divyanshu Kakwani , Dong Ahn , Duncan Riach , Dusan Stosic , Edgar Minasyan , Edward Lin , Eileen Long , Eileen Peters Long , Elena Lantz , Ellie Evans , Elliott Ning , Eric Chung , Eric Harper , Eric Tramel , Erick Galinkin , Erik Pounds , Evan Briones , Evelina Bakhturina , Faisal Ladhak , Fay Wang , Fei Jia , Felipe Soares , Feng Chen , Ferenc Galko , Frankie Siino , Gal Hubara Agam , Ganesh Ajjanagadde , Gantavya Bhatt , Gargi Prasad , George Armstrong , Gerald Shen , Gorkem Batmaz , Grigor Nalbandyan , Haifeng Qian , Harsh Sharma , Hayley Ross , Helen Ngo , Herman Sahota , Hexin Wang , Himanshu Soni , Hiren Upadhyay , Huizi Mao , Huy C Nguyen , Huy Q Nguyen , Iain Cunningham , Ido Shahaf , Igor Gitman , Ilya Loshchilov , Ivan Moshkov , Izzy Putterman , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jatin Mitra , Jeffrey Glick , Jenny Chen , Jesse Oliver , Jian Zhang , Jiaqi Zeng , Jie Lou , Jimmy Zhang , Jining Huang , Joey Conway , Joey Guman , John Kamalu , Johnny Greco , Jonathan Cohen , Joseph Jennings , Joyjit Daw , Julien Veron Vialard , Junkeun Yi , Jupinder Parmar , Kai Xu , Kan Zhu , Kari Briski , Katherine Cheung , Katherine Luna , Keshav Santhanam , Kevin Shih , Kezhi Kong , Khushi Bhardwaj , Krishna C. Puvvada , Krzysztof Pawelec , Kumar Anik , Lawrence McAfee , Laya Sleiman , Leon Derczynski , Li Ding , Lucas Liebenwein , Luis Vega , Maanu Grover , Maarten Van Segbroeck , Maer Rodrigues de Melo , Makesh Narsimhan Sreedhar , Manoj Kilaru , Maor Ashkenazi , Marc Romeijn , Mark Cai , Markus Kliegl , Maryam Moosaei , Matvei Novikov , Mehrzad Samadi , Melissa Corpuz , Mengru Wang , Meredith Price , Michael Boone , Michael Evans , Miguel Martinez , Mike Chrzanowski , Mohammad Shoeybi , Mostofa Patwary , Nabin Mulepati , Natalie Hereth , Nave Assaf , Negar Habibi , Neta Zmora , Netanel Haber , Nicola Sessions , Nidhi Bhatia , Nikhil Jukar , Nikki Pope , Nikolai Ludwig , Nima Tajbakhsh , Nirmal Juluru , Oleksii Hrinchuk , Oleksii Kuchaiev , Olivier Delalleau , Oluwatobi Olabiyi , Omer Ullman Argov , Ouye Xie , Parth Chadha , Pasha Shamis , Pavlo Molchanov , Pawel Morkisz , Peter Dykas , Peter Jin , Pinky Xu , Piotr Januszewski , Pranav Prashant Thombre , Prasoon Varshney , Pritam Gundecha , Qing Miao , Rabeeh Karimi Mahabadi , Ran El-Yaniv , Ran Zilberstein , Rasoul Shafipour , Rich Harang , Rick Izzo , Rima Shahbazyan , Rishabh Garg , Ritika Borkar , Ritu Gala , Riyad Islam , Roger Waleffe , Rohit Watve , Roi Koren , Ruoxi Zhang , Russell J. Hewett , Ryan Prenger , Ryan Timbrook , Sadegh Mahdavi , Sahil Modi , Samuel Kriman , Sanjay Kariyappa , Sanjeev Satheesh , Saori Kaji , Satish Pasumarthi , Sean Narentharen , Sean Narenthiran , Seonmyeong Bak , Sergey Kashirsky , Seth Poulos , Shahar Mor , Shanmugam Ramasamy , Shantanu Acharya , Shaona Ghosh , Sharath Turuvekere Sreenivas , Shelby Thomas , Shiqing Fan , Shreya Gopal , Shrimai Prabhumoye , Shubham Pachori , Shubham Toshniwal , Shuoyang Ding , Siddharth Singh , Simeng Sun , Smita Ithape , Somshubra Majumdar , Soumye Singhal , Stefania Alborghetti , Stephen Ge , Sugam Dipak Devare , Sumeet Kumar Barua , Suseella Panguluri , Suyog Gupta , Sweta Priyadarshi , Syeda Nahida Akter , Tan Bui , Teodor-Dumitru Ene , Terry Kong , Thanh Do , Tijmen Blankevoort , Tom Balough , Tomer Asida , Tomer Bar Natan , Tugrul Konuk , Twinkle Vashishth , Udi Karpas , Ushnish De , Vahid Noorozi , Vahid Noroozi , Venkat Srinivasan , Venmugil Elango , Vijay Korthikanti , Vitaly Kurin , Vitaly Lavrukhin , Wanli Jiang , Wasi Uddin Ahmad , Wei Du , Wei Ping , Wenfei Zhou , Will Jennings , William Zhang , Wojciech Prazuch , Xiaowei Ren , Yashaswi Karnati , Yejin Choi , Yev Meyer , Yi-Fu Wu , Yian Zhang , Ying Lin , Yonatan Geifman , Yonggan Fu , Yoshi Subara , Yoshi Suhara , Yubo Gao , Zach Moshe , Zhen Dong , Zihan Liu , Zijia Chen , Zijie Yan

Understanding videos is one of the fundamental directions in computer vision research, with extensive efforts dedicated to exploring various architectures such as RNN, 3D CNN, and Transformers. The newly proposed architecture of state space…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Guo Chen , Yifei Huang , Jilan Xu , Baoqi Pei , Zhe Chen , Zhiqi Li , Jiahao Wang , Kunchang Li , Tong Lu , Limin Wang

Large language models (LLMs) have shown continuously improving multilingual capabilities, and even small-scale open-source models have demonstrated rapid performance enhancement. In this paper, we systematically explore the abilities of…

Computation and Language · Computer Science 2025-02-25 Menglong Cui , Pengzhi Gao , Wei Liu , Jian Luan , Bin Wang

We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has $225.8$B total parameters ($23.4$B activated per token) and XS.2 has $33.4$B total ($3$B activated). Both models…

Artificial Intelligence · Computer Science 2026-05-28 Julien Abadji , Marah Abdin , Connor Adams , Eric Alcaide , Mustafa Altun , Michele Artoni , Junze Bao , Uday Barar , Vassilis Bekiaris , Arkadii Bessonov , Benjamin Bütikofer , Jonathan Chang , Yen-Chun Chen , Dmitry Chernenkov , Yang Chi , Filippos Christianos , Fenia Christopoulou , Razvan-Andrei Ciocoiu , Tzachi Cohen , Yohann Coppel , Dmitrii Emelianenko , Brandon Fergerson , Brian Fitzgerald , Matthias Gallé , Alex Golonzovskyi , George Grigorev , Yiyang Hao , Christian Hensel , Jan Huenermann , Ye Ji , Sarthak Joshi , Eiso Kant , Kabir Khandpur , Seonghyeon Kim , Vladimir Kirichenko , Umut Kocasarac , Ilya Kochik , Ivan Komarov , Chaerin Kong , Anurag Koul , François-Joseph Lacroix , Sergei Laktionov , Waren Long , Quentin Malartic , Vadim Markovtsev , Afonso Marques , Robert McHardy , Carlos Mocholí , Dmitry Monakhov , Adam Morris , Martin Muller , Christian Mürtz , Robin Nabel , Thien Nguyen , Rok Novosel , Szymon Ozog , Aalhad Patankar , Aleksei Petrov , Alexandre Piché , Arthur Pignet , Teodor Poncu , Phil Potter , Alexander Rakowski , Pierre-Yves Ritschard , Jay Roberts , Joe Rowell , Piotr Sarna , Pierre-André Savalle , Uladzislau Sazanovich , Nikita Shapovalov , Arsenii Shevchenko , Mikhail Shilkov , Andrei Sokol , Mohamed Soliman , Jack Stephenson , Victor Storchan , Dragos-Constantin Tantaru , Artem Tyurin , Adrian Wälchli , Pengming Wang , Jianxiao Yang , Renat Zayashnikov , Alexander Zelenka Martin , Nikolay Zinov , Caroline Bercier , José Caldeira , Margarida Garcia , Tom George , Kabeer Gharzai , Glenn Hitchcock , Carson Klingenberg , Ivo Pinto , Varun Randery , Noah Smith , Arina Sugako , Jason Warner

Yuan 2.0-M32, with a similar base architecture as Yuan-2.0 2B, uses a mixture-of-experts architecture with 32 experts of which 2 experts are active. A new router network, Attention Router, is proposed and adopted for a more efficient…

Artificial Intelligence · Computer Science 2024-05-30 Shaohua Wu , Jiangang Luo , Xi Chen , Lingjun Li , Xudong Zhao , Tong Yu , Chao Wang , Yue Wang , Fei Wang , Weixu Qiao , Houbo He , Zeru Zhang , Zeyu Sun , Junxiong Mao , Chong Shen

Breeze-7B is an open-source language model based on Mistral-7B, designed to address the need for improved language comprehension and chatbot-oriented capabilities in Traditional Chinese. This technical report provides an overview of the…

Computation and Language · Computer Science 2024-04-04 Chan-Jan Hsu , Chang-Le Liu , Feng-Ting Liao , Po-Chun Hsu , Yi-Chang Chen , Da-Shan Shiu

Large language models operate in distinct compute-bound prefill followed by memory bandwidth-bound decode phases. Hybrid Mamba-Transformer models inherit this asymmetry while adding state space model (SSM) recurrences and element-wise…

Hardware Architecture · Computer Science 2026-03-17 Alish Kanani , Sangwan Lee , Han Lyu , Jiahao Lin , Jaehyun Park , Umit Y. Ogras

EngGPT2-16B-A3B is the latest iteration of Engineering Group's Italian LLM and it's built to be a Sovereign, Efficient and Open model. EngGPT2 is trained on 2.5 trillion tokens - less than Qwen3's 36T or Llama3's 15T - and delivers…

This paper starts with a simple lossless ~1.5:1 compression algorithm for the weights of the Large Language Model (LLM) Llama2 7B [1] that can be implemented in ~200 LUTs in AMD FPGAs, processing over 800 million bfloat16 numbers per…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Vincenzo Liguori

State-space models (SSMs) have recently demonstrated competitive performance to transformers at large-scale language modeling benchmarks while achieving linear time and memory complexity as a function of sequence length. Mamba, a recently…

Computation and Language · Computer Science 2024-02-06 Quentin Anthony , Yury Tokpanov , Paolo Glorioso , Beren Millidge

In recent years, the talking head generation has become a focal point for researchers. Considerable effort is being made to refine lip-sync motion, capture expressive facial expressions, generate natural head poses, and achieve high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Farzaneh Jafari , Stefano Berretti , Anup Basu

We introduce Motif-2-12.7B, a new open-weight foundation model that pushes the efficiency frontier of large language models by combining architectural innovation with system-level optimization. Designed for scalable language understanding…

We present Sapiens2, a model family of high-resolution transformers for human-centric vision focused on generalization, versatility, and high-fidelity outputs. Our model sizes range from 0.4 to 5 billion parameters, with native 1K…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Rawal Khirodkar , He Wen , Julieta Martinez , Yuan Dong , Su Zhaoen , Shunsuke Saito
‹ Prev 1 4 5 6 7 8 10 Next ›