English
Related papers

Related papers: Nemotron Elastic: Towards Efficient Many-in-One Re…

200 papers

Reinforcement Learning (RL) has shown promise in improving the reasoning abilities of Large Language Models (LLMs). However, the specific challenges of adapting RL to multimodal data and formats remain relatively unexplored. In this work,…

Machine Learning · Computer Science 2025-05-20 Zirun Guo , Minjie Hong , Tao Jin

Being prompted to engage in reasoning has emerged as a core technique for using large language models (LLMs), deploying additional inference-time compute to improve task performance. However, as LLMs increase in both size and adoption,…

Computation and Language · Computer Science 2025-06-25 C. Nicolò De Sabbata , Theodore R. Sumers , Badr AlKhamissi , Antoine Bosselut , Thomas L. Griffiths

Large Reasoning Models (LRMs) often suffer from overthinking, generating unnecessarily long reasoning chains even for simple tasks. This leads to substantial computational overhead with limited performance gain, primarily due to redundant…

Artificial Intelligence · Computer Science 2026-01-13 Ruichu Cai , Haopeng Du , Qingwen Lin , Yutong Chen , Zijian Li , Boyan Xu

The considerable size of Large Language Models (LLMs) presents notable deployment challenges, particularly on resource-constrained hardware. Structured pruning, offers an effective means to compress LLMs, thereby reducing storage costs and…

Computation and Language · Computer Science 2024-06-28 Shengrui Li , Junzhe Chen , Xueting Han , Jing Bai

Large language models (LLMs) have demonstrated impressive capabilities across various tasks, but the billion-scale parameters pose deployment challenges. Although existing methods attempt to reduce the scale of LLMs, they require either…

Computation and Language · Computer Science 2026-04-07 Xinhao Huang , You-Liang Huang , Zeyi Wen

The recent surge in Multimodal Large Language Models (MLLMs) has showcased their remarkable potential for achieving generalized intelligence by integrating visual understanding into Large Language Models.Nevertheless, the sheer model size…

Computation and Language · Computer Science 2024-07-30 Shilin Xu , Xiangtai Li , Haobo Yuan , Lu Qi , Yunhai Tong , Ming-Hsuan Yang

Multimodal Large Language Models (MLLMs) have demonstrated exceptional success in various multimodal tasks, yet their deployment is frequently limited by substantial computational demands and prolonged inference times. Given that the vision…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Zihui Zhao , Yingxin Li , Yang Li

We study distillation for large language models under explicit compute constraints, with the goal of producing student models that are not only cheaper to train, but structurally efficient at inference time. While prior approaches to…

Machine Learning · Computer Science 2026-05-07 Mohammed Sabry , Anya Belz

With the growing demand for deploying large language models (LLMs) across diverse applications, improving their inference efficiency is crucial for sustainable and democratized access. However, retraining LLMs to meet new user-specific…

Machine Learning · Computer Science 2026-01-21 Mingyu Yang , Mehdi Rezagholizadeh , Guihong Li , Vikram Appia , Emad Barsoum

Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains, mathematical reasoning serves as a representative benchmark…

Large language models (LLMs) have exhibited impressive reasoning abilities on a wide range of complex tasks. However, enhancing these capabilities through post-training remains resource intensive, particularly in terms of data and…

Artificial Intelligence · Computer Science 2025-08-13 Shuo Cai , Su Lu , Qi Zhou , Kejing Yang , Zhijie Sang , Congkai Xie , Hongxia Yang

The scalability of large language models for long-context reasoning is severely constrained by the linear growth of their Transformer key-value cache, which incurs significant memory and computational costs. We posit that as a model…

Computation and Language · Computer Science 2025-12-30 Giovanni Monea , Yair Feldman , Shankar Padmanabhan , Kianté Brantley , Yoav Artzi

Large Language Models (LLMs) using Chain-of-Thought (CoT) prompting excel at complex reasoning but generate verbose thought processes with considerable redundancy, leading to increased inference costs and reduced efficiency. We introduce a…

Artificial Intelligence · Computer Science 2026-02-17 Zeju Li , Jianyuan Zhong , Ziyang Zheng , Xiangyu Wen , Zhijian Xu , Yingying Cheng , Fan Zhang , Qiang Xu

Despite the significant improvements achieved by large language models (LLMs) in English reasoning tasks, these models continue to struggle with multilingual reasoning. Recent studies leverage a full-parameter and two-stage training…

Computation and Language · Computer Science 2025-01-08 Yuchun Fan , Yongyu Mu , Yilin Wang , Lei Huang , Junhao Ruan , Bei Li , Tong Xiao , Shujian Huang , Xiaocheng Feng , Jingbo Zhu

Large Language Models (LLMs) achieve superior performance through Chain-of-Thought (CoT) reasoning, but these token-level reasoning chains are computationally expensive and inefficient. In this paper, we introduce Compressed Latent…

Computation and Language · Computer Science 2026-02-04 Wenhui Tan , Jiaze Li , Jianzhong Ju , Zhenbo Luo , Ruihua Song , Jian Luan

We introduce Nemotron Nano V2 VL, the latest model of the Nemotron vision-language series designed for strong real-world document understanding, long video comprehension, and reasoning tasks. Nemotron Nano V2 VL delivers significant…

Machine Learning · Computer Science 2025-11-10 NVIDIA , : , Amala Sanjay Deshmukh , Kateryna Chumachenko , Tuomas Rintamaki , Matthieu Le , Tyler Poon , Danial Mohseni Taheri , Ilia Karmanov , Guilin Liu , Jarno Seppanen , Guo Chen , Karan Sapra , Zhiding Yu , Adi Renduchintala , Charles Wang , Peter Jin , Arushi Goel , Mike Ranzinger , Lukas Voegtle , Philipp Fischer , Timo Roman , Wei Ping , Boxin Wang , Zhuolin Yang , Nayeon Lee , Shaokun Zhang , Fuxiao Liu , Zhiqi Li , Di Zhang , Greg Heinrich , Hongxu Yin , Song Han , Pavlo Molchanov , Parth Mannan , Yao Xu , Jane Polak Scowcroft , Tom Balough , Subhashree Radhakrishnan , Paris Zhang , Sean Cha , Ratnesh Kumar , Zaid Pervaiz Bhat , Jian Zhang , Darragh Hanley , Pritam Biswas , Jesse Oliver , Kevin Vasques , Roger Waleffe , Duncan Riach , Oluwatobi Olabiyi , Ameya Sunil Mahabaleshwarkar , Bilal Kartal , Pritam Gundecha , Khanh Nguyen , Alexandre Milesi , Eugene Khvedchenia , Ran Zilberstein , Ofri Masad , Natan Bagrov , Nave Assaf , Tomer Asida , Daniel Afrimi , Amit Zuker , Netanel Haber , Zhiyu Cheng , Jingyu Xin , Di Wu , Nik Spirin , Maryam Moosaei , Roman Ageev , Vanshil Atul Shah , Yuting Wu , Daniel Korzekwa , Unnikrishnan Kizhakkemadam Sreekumar , Wanli Jiang , Padmavathy Subramanian , Alejandra Rico , Sandip Bhaskar , Saeid Motiian , Kedi Wu , Annie Surla , Chia-Chih Chen , Hayden Wolff , Matthew Feinberg , Melissa Corpuz , Marek Wawrzos , Eileen Long , Aastha Jhunjhunwala , Paul Hendricks , Farzan Memarian , Benika Hall , Xin-Yu Wang , David Mosallanezhad , Soumye Singhal , Luis Vega , Katherine Cheung , Krzysztof Pawelec , Michael Evans , Katherine Luna , Jie Lou , Erick Galinkin , Akshay Hazare , Kaustubh Purandare , Ann Guan , Anna Warno , Chen Cui , Yoshi Suhara , Shibani Likhite , Seph Mard , Meredith Price , Laya Sleiman , Saori Kaji , Udi Karpas , Kari Briski , Joey Conway , Michael Lightstone , Jan Kautz , Mohammad Shoeybi , Mostofa Patwary , Jonathen Cohen , Oleksii Kuchaiev , Andrew Tao , Bryan Catanzaro

Transformer-based large language models (LLMs) are increasingly being adopted in networking research to address domain-specific challenges. However, their quadratic time complexity and substantial model sizes often result in significant…

Networking and Internet Architecture · Computer Science 2025-10-21 Linhan Xia , Mingzhan Yang , Jingjing Wang , Ziwei Yan , Yakun Ren , Guo Yu , Kai Lei

Recent breakthroughs in generative reasoning have fundamentally reshaped how large language models (LLMs) address complex tasks, enabling them to dynamically retrieve, refine, and organize information into coherent multi-step reasoning…

Machine Learning · Computer Science 2026-01-06 Mohamed Amine Ferrag , Norbert Tihanyi , Merouane Debbah

Efficiently updating Large Language Models (LLMs) with new or evolving factual knowledge remains a central challenge, as even parameter-efficient adaptation can erode previously acquired reasoning abilities. This tension reflects a…

Artificial Intelligence · Computer Science 2026-05-26 Mustafa Hayri Bilgin , Mariam Barry , Albert Bifet , Azzedine Idir Ait Said , Soumya Banerjee

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture,…