English
Related papers

Related papers: Llama-Nemotron: Efficient Reasoning Models

200 papers

Recent advancements in large language models (LLMs), such as DeepSeek-R1 and OpenAI-o1, have demonstrated the significant effectiveness of test-time scaling, achieving substantial performance gains across various benchmarks. These advanced…

Computation and Language · Computer Science 2025-04-15 Haotian Wang , Han Zhao , Shuaiting Chen , Xiaoyu Tian , Sitong Zhao , Yunjie Ji , Yiping Peng , Xiangang Li

We propose cognitive prompting as a novel approach to guide problem-solving in large language models (LLMs) through structured, human-like cognitive operations, such as goal clarification, decomposition, filtering, abstraction, and pattern…

Computation and Language · Computer Science 2024-12-03 Oliver Kramer , Jill Baumann

Large language models (LLMs) tackle complex tasks by generating long chains of thought or "reasoning traces" that act as latent variables in the generation of an output given a query. A model's ability to generate such traces can be…

Computation and Language · Computer Science 2025-12-03 Alexander Gurung , Nikolay Malkin , Mirella Lapata

Large Language Models (LLMs) with reasoning capabilities have achieved state-of-the-art performance on a wide range of tasks. Despite its empirical success, the tasks and model scales at which reasoning becomes effective, as well as its…

Computation and Language · Computer Science 2025-09-29 Nicolas Boizard , Hippolyte Gisserot-Boukhlef , Kevin El-Haddad , Céline Hudelot , Pierre Colombo

Neural-symbolic methods have demonstrated efficiency in enhancing the reasoning abilities of large language models (LLMs). However, existing methods mainly rely on syntactically mapping natural languages to complete formal languages like…

Computation and Language · Computer Science 2024-06-04 Yiming Wang , Zhuosheng Zhang , Pei Zhang , Baosong Yang , Rui Wang

Theory of Mind (ToM) assesses whether models can infer hidden mental states such as beliefs, desires, and intentions, which is essential for natural social interaction. Although recent progress in Large Reasoning Models (LRMs) has boosted…

Artificial Intelligence · Computer Science 2026-03-05 Nanxu Gong , Haotian Li , Sixun Dong , Jianxun Lian , Yanjie Fu , Xing Xie

We introduce FFN Fusion, an architectural optimization technique that reduces sequential computation in large language models by identifying and exploiting natural opportunities for parallelization. Our key insight is that sequences of…

Prompting techniques such as chain-of-thought have established themselves as a popular vehicle for improving the outputs of large language models (LLMs). For code generation, however, their exact mechanics and efficacy are under-explored.…

Computation and Language · Computer Science 2025-04-09 Kunhao Zheng , Juliette Decugis , Jonas Gehring , Taco Cohen , Benjamin Negrevergne , Gabriel Synnaeve

Large language models (LLMs) can potentially democratize access to medical knowledge. While many efforts have been made to harness and improve LLMs' medical knowledge and reasoning capacities, the resulting models are either closed-source…

Recent advancements in Large Language Models (LLMs) have significantly enhanced their ability to perform complex reasoning tasks, transitioning from fast and intuitive thinking (System 1) to slow and deep reasoning (System 2). While System…

Computation and Language · Computer Science 2025-04-01 Rui Wang , Hongru Wang , Boyang Xue , Jianhui Pang , Shudong Liu , Yi Chen , Jiahao Qiu , Derek Fai Wong , Heng Ji , Kam-Fai Wong

Recent advances in Multi-Modal Large Language Models (MLLMs) have enabled unified processing of language, vision, and structured inputs, opening the door to complex tasks such as logical deduction, spatial reasoning, and scientific…

Artificial Intelligence · Computer Science 2025-07-03 Guiyao Tie , Xueyang Zhou , Tianhe Gu , Ruihang Zhang , Chaoran Hu , Sizhe Zhang , Mengqu Sun , Yan Zhang , Pan Zhou , Lichao Sun

The math abilities of large language models can represent their abstract reasoning ability. In this paper, we introduce and open-source our math reasoning LLMs InternLM-Math which is continue pre-trained from InternLM2. We unify…

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its…

Machine Learning · Computer Science 2026-05-12 NVIDIA , : , Amala Sanjay Deshmukh , Kateryna Chumachenko , Tuomas Rintamaki , Matthieu Le , Tyler Poon , Danial Mohseni Taheri , Ilia Karmanov , Guilin Liu , Jarno Seppanen , Arushi Goel , Mike Ranzinger , Greg Heinrich , Guo Chen , Lukas Voegtle , Philipp Fischer , Timo Roman , Karan Sapra , Collin McCarthy , Shaokun Zhang , Fuxiao Liu , Hanrong Ye , Yi Dong , Mingjie Liu , Yifan Peng , Piotr Zelasko , Zhehuai Chen , Nithin Rao Koluguri , Nune Tadevosyan , Lilit Grigoryan , Ehsan Hosseini Asl , Pritam Biswas , Leili Tavabi , Yuanhang Su , Zhiding Yu , Peter Jin , Alexandre Milesi , Netanel Haber , Yao Xu , Sarah Amiraslani , Nabin Mulepati , Eric Tramel , Jaehun Jung , Ximing Lu , Brandon Cui , Jin Xu , Zhiqi Li , Shihao Wang , Yuanguo Kuang , Shaokun Zhang , Huck Yang , Boyi Li , Hongxu Yin , Song Han , Bilal Kartal , Pavlo Molchanov , Adi Renduchintala , Charles Wang , David Mosallanezhad , Soumye Singhal , Luis Vega , Katherine Cheung , Sreyan Ghosh , Yian Zhang , Alexander Bukharin , Venkat Srinivasan , Johnny Greco , Andre Manoel , Maarten Van Segbroeck , Suseella Panguliri , Rohit Watve , Divyanshu Kakwani , Shubham Pachori , Jeffrey Glick , Radha Sri-Tharan , Aileen Zaman , Khanh Nguyen , Shi Chen , Jiaheng Fang , Qing Miao , Wenfei Zhou , Yu Wang , Zaid Pervaiz Bhat , Varun Praveen , Arihant Jain , Ramanathan Arunachalam , Tomasz Kornuta , Ashton Sharabiani , Amy Shen , Wei Huang , Yi-Fu Wu , Ali Roshan Ghias , Huiying Li , Brian Yu , Nima Tajbakhsh , Chen Cui , Wenwen Gao , Li Ding , Terry Kong , Manoj Kilaru , Anahita Bhiwandiwalla , Marek Wawrzos , Daniel Korzekwa , Pablo Ribalta , Grzegorz Chlebus , Besmira Nushi , Ewa Dobrowolska , Maciej Jakub Mikulski , Kunal Dhawan , Steve Huang , Jagadeesh Balam , Yongqiang Wang , Nikolay Karpov , Valentin Mendelev , George Zelenfroynd , Meline Mkrtchyan , Qing Miao , Omri Almog , Bhavesh Pawar , Rameshwar Shivbhakta , Sudeep Sabnis , Ashrton Sharabiani , Negar Habibi , Geethapriya Venkataramani , Pamela Peng , Prerit Rodney , Serge Panev , Richard Mazzarese , Nicky Liu , Michael Fukuyama , Andrii Skliar , Roger Waleffe , Duncan Riach , Yunheng Zou , Jian Hu , Hao Zhang , Binfeng Xu , Yuhao Yang , Zuhair Ahmed , Alexandre Milesi , Carlo del Mundo , Chad Voegele , Zhiyu Cheng , Nave Assaf , Andrii Skliar , Daniel Afrimi , Natan Bagrov , Ran Zilberstein , Ofri Masad , Eugene Khvedchenia , Natan Bagrov , Borys Tymchenko , Tomer Asida , Daniel Afrimi , Parth Mannan , Victor Cui , Michael Evans , Katherine Luna , Jie Lou , Pinky Xu , Guyue Huang , Negar Habibi , Michael Boone , Pradeep Thalasta , Adeola Adesoba , Dina Yared , Christopher Parisien , Leon Derczynski , Shaona Ghosh , Wes Feely , Micah Schaffer , Radha Sri-Tharan , Jeffrey Glick , Barnaby Simkin , George Zelenfroynd , Tomasz Grzegorzek , Rishabh Garg , Aastha Jhunjhunwala , Sergei Kolchenko , Farzan Memarian , Haran Kumar , Shiv Kumar , Isabel Hulseman , Anjali Shah , Kari Briski , Padmavathy Subramanian , Joey Conway , Udi Karpas , Jane Polak Scowcroft , Annie Surla , Shilpa Ammireddy , Ellie Evans , Jesse Oliver , Tom Balough , Chia-Chih Chen , Sandip Bhaskar , Alejandra Rico , Bardiya Sadeghi , Seph Mard , Katherine Cheung , Meredith Price , Laya Sleiman , Saori Kaji , Wesley Helmholz , Wendy Quan , Michael Lightstone , Jonathan Cohen , Jian Zhang , Oleksii Kuchaiev , Boris Ginsburg , Jan Kautz , Eileen Long , Mohammad Shoeybi , Mostofa Patwary , Oluwatobi Olabiyi , Andrew Tao , Bryan Catanzaro , Udi Karpas

While large language models (LLMs) excel in mathematical and code reasoning, we observe they struggle with social reasoning tasks, exhibiting cognitive confusion, logical inconsistencies, and conflation between objective world states and…

Computation and Language · Computer Science 2025-10-14 Jialu Du , Guiyang Hou , Yihui Fu , Chen Wu , Wenqi Zhang , Yongliang Shen , Weiming Lu

Large language models (LLMs) are useful in many NLP tasks and become more capable with size, with the best open-source models having over 50 billion parameters. However, using these 50B+ models requires high-end hardware, making them…

Human cognition is theorized to operate in two modes: fast, intuitive System 1 thinking and slow, deliberate System 2 thinking. While current Large Reasoning Models (LRMs) excel at System 2 thinking, their inability to perform fast thinking…

Computation and Language · Computer Science 2025-10-31 Zhengkai Lin , Zhihang Fu , Ze Chen , Chao Chen , Liang Xie , Wenxiao Wang , Deng Cai , Zheng Wang , Jieping Ye

Reasoning Language Models, capable of extended chain-of-thought reasoning, have demonstrated remarkable performance on tasks requiring complex logical inference. However, applying elaborate reasoning for all queries often results in…

Computation and Language · Computer Science 2025-06-27 Gongfan Fang , Xinyin Ma , Xinchao Wang

Recent advances in fine-tuning large language models (LLMs) with reinforcement learning (RL) have shown promising improvements in complex reasoning tasks, particularly when paired with chain-of-thought (CoT) prompting. However, these…

Machine Learning · Computer Science 2025-04-04 Hung Le , Dai Do , Dung Nguyen , Svetha Venkatesh

Compressing long chain-of-thought (CoT) from large language models (LLMs) is an emerging strategy to improve the reasoning efficiency of LLMs. Despite its promising benefits, existing studies equally compress all thoughts within a long CoT,…

Computation and Language · Computer Science 2025-05-27 Yansong Ning , Wei Li , Jun Fang , Naiqiang Tan , Hao Liu

Large Reasoning Models (LRMs) have achieved remarkable performance on complex reasoning tasks by adopting the ``think-then-answer'' paradigm, which enhances both accuracy and interpretability. However, current LRMs exhibit two critical…

Computation and Language · Computer Science 2026-01-09 Xue Zhang , Yunlong Liang , Fandong Meng , Songming Zhang , Kaiyu Huang , Yufeng Chen , Jinan Xu , Jie Zhou
‹ Prev 1 3 4 5 6 7 10 Next ›