English
Related papers

Related papers: The Zamba2 Suite: Technical Report

200 papers

Transformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with quadratic complexity is inefficient in computation and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-02 Xilin Jiang , Cong Han , Nima Mesgarani

We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert transformer architecture and are trained using large-scale…

Computation and Language · Computer Science 2025-08-18 OpenAI , : , Sandhini Agarwal , Lama Ahmad , Jason Ai , Sam Altman , Andy Applebaum , Edwin Arbus , Rahul K. Arora , Yu Bai , Bowen Baker , Haiming Bao , Boaz Barak , Ally Bennett , Tyler Bertao , Nivedita Brett , Eugene Brevdo , Greg Brockman , Sebastien Bubeck , Che Chang , Kai Chen , Mark Chen , Enoch Cheung , Aidan Clark , Dan Cook , Marat Dukhan , Casey Dvorak , Kevin Fives , Vlad Fomenko , Timur Garipov , Kristian Georgiev , Mia Glaese , Tarun Gogineni , Adam Goucher , Lukas Gross , Katia Gil Guzman , John Hallman , Jackie Hehir , Johannes Heidecke , Alec Helyar , Haitang Hu , Romain Huet , Jacob Huh , Saachi Jain , Zach Johnson , Chris Koch , Irina Kofman , Dominik Kundel , Jason Kwon , Volodymyr Kyrylov , Elaine Ya Le , Guillaume Leclerc , James Park Lennon , Scott Lessans , Mario Lezcano-Casado , Yuanzhi Li , Zhuohan Li , Ji Lin , Jordan Liss , Lily , Liu , Jiancheng Liu , Kevin Lu , Chris Lu , Zoran Martinovic , Lindsay McCallum , Josh McGrath , Scott McKinney , Aidan McLaughlin , Song Mei , Steve Mostovoy , Tong Mu , Gideon Myles , Alexander Neitz , Alex Nichol , Jakub Pachocki , Alex Paino , Dana Palmie , Ashley Pantuliano , Giambattista Parascandolo , Jongsoo Park , Leher Pathak , Carolina Paz , Ludovic Peran , Dmitry Pimenov , Michelle Pokrass , Elizabeth Proehl , Huida Qiu , Gaby Raila , Filippo Raso , Hongyu Ren , Kimmy Richardson , David Robinson , Bob Rotsted , Hadi Salman , Suvansh Sanjeev , Max Schwarzer , D. Sculley , Harshit Sikchi , Kendal Simon , Karan Singhal , Yang Song , Dane Stuckey , Zhiqing Sun , Philippe Tillet , Sam Toizer , Foivos Tsimpourlas , Nikhil Vyas , Eric Wallace , Xin Wang , Miles Wang , Olivia Watkins , Kevin Weil , Amy Wendling , Kevin Whinnery , Cedric Whitney , Hannah Wong , Lin Yang , Yu Yang , Michihiro Yasunaga , Kristen Ying , Wojciech Zaremba , Wenting Zhan , Cyril Zhang , Brian Zhang , Eddie Zhang , Shengjia Zhao

Recent advancements in imitation learning have been largely fueled by the integration of sequence models, which provide a structured flow of information to effectively mimic task behaviours. Currently, Decision Transformer (DT) and…

Machine Learning · Computer Science 2024-10-18 André Correia , Luís A. Alexandre

Multi-view depth estimation has achieved impressive performance over various benchmarks. However, almost all current multi-view systems rely on given ideal camera poses, which are unavailable in many real-world scenarios, such as autonomous…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Zelin Meng , Zhichen Wang

The quadratic computational complexity of self-attention in diffusion transformers (DiT) introduces substantial computational costs in high-resolution image generation. While the linear-complexity Mamba model emerges as a potential…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Yuan Yao , Yicong Hong , Difan Liu , Long Mai , Feng Liu , Jiebo Luo

We introduce Falcon2-11B, a foundation model trained on over five trillion tokens, and its multimodal counterpart, Falcon2-11B-vlm, which is a vision-to-text model. We report our findings during the training of the Falcon2-11B which follows…

Xmodel-2 is a 1.2-billion-parameter large language model designed specifically for reasoning tasks. Its architecture enables different model scales to share a unified set of hyperparameters, allowing for extensive experimentation on smaller…

Artificial Intelligence · Computer Science 2024-12-30 Wang Qun , Liu Yang , Lin Qingquan , Qu Zhijiu , Jiang Ling

Recent advancements in State Space Models (SSMs) have attracted significant interest, particularly in models optimized for parallel training and handling long-range dependencies. Architectures like Mamba have scaled to billions of…

Machine Learning · Computer Science 2024-10-22 Zheng Zhan , Yushu Wu , Zhenglun Kong , Changdi Yang , Yifan Gong , Xuan Shen , Xue Lin , Pu Zhao , Yanzhi Wang

Large language models (LLMs) have showcased profound capabilities in language understanding and generation, facilitating a wide array of applications. However, there is a notable paucity of detailed, open-sourced methodologies on…

Time series forecasting has made significant advances, including with Transformer-based models. The attention mechanism in Transformer effectively captures temporal dependencies by attending to all past inputs simultaneously. However, its…

Machine Learning · Computer Science 2025-11-04 Xiongxiao Xu , Canyu Chen , Yueqing Liang , Baixiang Huang , Guangji Bai , Liang Zhao , Kai Shu

Mamba is an emerging, complex workload with various short-range and long-range dependencies, nonlinearities, and elementwise computations that are unable to run at near-peak speeds on modern hardware. Specifically, Mamba's complex…

Hardware Architecture · Computer Science 2026-04-07 Toluwanimi O. Odemuyiwa , John D. Owens , Joel S. Emer , Michael Pellauer

Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks, yet the majority of high-performing models remain closed-source or partially open, limiting transparency and reproducibility. In this work,…

Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead. We…

Machine Learning · Computer Science 2026-03-02 Vaibhav Singh , Oleksiy Ostapenko , Pierre-André Noël , Eugene Belilovsky , Torsten Scholak

We present ZAYA1-VL-8B, a compact mixture-of-experts vision-language model built upon our in-house language model, ZAYA1-8B. Despite its compact size, ZAYA1-VL achieves performance competitive with leading base models such as Molmo2-4B and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Hassan Shapourian , Kasra Hejazi , Olabode M. Sule , Beren Millidge

Transformers have become one of the foundational architectures in point cloud analysis tasks due to their excellent global modeling ability. However, the attention mechanism has quadratic complexity, making the design of a linear complexity…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Dingkang Liang , Xin Zhou , Wei Xu , Xingkui Zhu , Zhikang Zou , Xiaoqing Ye , Xiao Tan , Xiang Bai

Large language models (LLMs) have demonstrated remarkable performance on a variety of natural language tasks based on just a few examples of natural language instructions, reducing the need for extensive feature engineering. However, most…

In recent years, robust matching methods using deep learning-based approaches have been actively studied and improved in computer vision tasks. However, there remains a persistent demand for both robust and fast matching techniques. To…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Kihwan Ryoo , Hyungtae Lim , Hyun Myung

Large Language Models (LLMs) have achieved remarkable results, but their increasing resource demand has become a major obstacle to the development of powerful and accessible super-human intelligence. This report introduces JetMoE-8B, a new…

Computation and Language · Computer Science 2024-04-12 Yikang Shen , Zhen Guo , Tianle Cai , Zengyi Qin

In the realm of time series forecasting (TSF), it is imperative for models to adeptly discern and distill hidden patterns within historical time series data to forecast future states. Transformer-based models exhibit formidable efficacy in…

Machine Learning · Computer Science 2024-04-30 Zihan Wang , Fanheng Kong , Shi Feng , Ming Wang , Xiaocui Yang , Han Zhao , Daling Wang , Yifei Zhang

This paper examines the mathematical foundations of transformer architectures, highlighting their limitations particularly in handling long sequences. We explore prerequisite models such as Mamba, Vision Mamba (ViM), and LV-ViT that pave…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Ricky Fang