English
Related papers

Related papers: Accelerating PayPal's Commerce Agent with Speculat…

200 papers

We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to…

Computation and Language · Computer Science 2025-12-25 NVIDIA , : , Aaron Blakeman , Aaron Grattafiori , Aarti Basant , Abhibha Gupta , Abhinav Khattar , Adi Renduchintala , Aditya Vavre , Akanksha Shukla , Akhiad Bercovich , Aleksander Ficek , Aleksandr Shaposhnikov , Alex Kondratenko , Alexander Bukharin , Alexandre Milesi , Ali Taghibakhshi , Alisa Liu , Amelia Barton , Ameya Sunil Mahabaleshwarkar , Amir Klein , Amit Zuker , Amnon Geifman , Amy Shen , Anahita Bhiwandiwalla , Andrew Tao , Anjulie Agrusa , Ankur Verma , Ann Guan , Anubhav Mandarwal , Arham Mehta , Ashwath Aithal , Ashwin Poojary , Asif Ahamed , Asit Mishra , Asma Kuriparambil Thekkumpate , Ayush Dattagupta , Banghua Zhu , Bardiya Sadeghi , Barnaby Simkin , Ben Lanir , Benedikt Schifferer , Besmira Nushi , Bilal Kartal , Bita Darvish Rouhani , Boris Ginsburg , Brandon Norick , Brandon Soubasis , Branislav Kisacanin , Brian Yu , Bryan Catanzaro , Carlo del Mundo , Chantal Hwang , Charles Wang , Cheng-Ping Hsieh , Chenghao Zhang , Chenhan Yu , Chetan Mungekar , Chintan Patel , Chris Alexiuk , Christopher Parisien , Collin Neale , Cyril Meurillon , Damon Mosk-Aoyama , Dan Su , Dane Corneil , Daniel Afrimi , Daniel Lo , Daniel Rohrer , Daniel Serebrenik , Daria Gitman , Daria Levy , Darko Stosic , David Mosallanezhad , Deepak Narayanan , Dhruv Nathawani , Dima Rekesh , Dina Yared , Divyanshu Kakwani , Dong Ahn , Duncan Riach , Dusan Stosic , Edgar Minasyan , Edward Lin , Eileen Long , Eileen Peters Long , Elad Segal , Elena Lantz , Ellie Evans , Elliott Ning , Eric Chung , Eric Harper , Eric Tramel , Erick Galinkin , Erik Pounds , Evan Briones , Evelina Bakhturina , Evgeny Tsykunov , Faisal Ladhak , Fay Wang , Fei Jia , Felipe Soares , Feng Chen , Ferenc Galko , Frank Sun , Frankie Siino , Gal Hubara Agam , Ganesh Ajjanagadde , Gantavya Bhatt , Gargi Prasad , George Armstrong , Gerald Shen , Gorkem Batmaz , Grigor Nalbandyan , Haifeng Qian , Harsh Sharma , Hayley Ross , Helen Ngo , Herbert Hum , Herman Sahota , Hexin Wang , Himanshu Soni , Hiren Upadhyay , Huizi Mao , Huy C Nguyen , Huy Q Nguyen , Iain Cunningham , Ido Galil , Ido Shahaf , Igor Gitman , Ilya Loshchilov , Itamar Schen , Itay Levy , Ivan Moshkov , Izik Golan , Izzy Putterman , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jatin Mitra , Jeffrey Glick , Jenny Chen , Jesse Oliver , Jian Zhang , Jiaqi Zeng , Jie Lou , Jimmy Zhang , Jinhang Choi , Jining Huang , Joey Conway , Joey Guman , John Kamalu , Johnny Greco , Jonathan Cohen , Joseph Jennings , Joyjit Daw , Julien Veron Vialard , Junkeun Yi , Jupinder Parmar , Kai Xu , Kan Zhu , Kari Briski , Katherine Cheung , Katherine Luna , Keith Wyss , Keshav Santhanam , Kevin Shih , Kezhi Kong , Khushi Bhardwaj , Kirthi Shankar , Krishna C. Puvvada , Krzysztof Pawelec , Kumar Anik , Lawrence McAfee , Laya Sleiman , Leon Derczynski , Li Ding , Lizzie Wei , Lucas Liebenwein , Luis Vega , Maanu Grover , Maarten Van Segbroeck , Maer Rodrigues de Melo , Mahdi Nazemi , Makesh Narsimhan Sreedhar , Manoj Kilaru , Maor Ashkenazi , Marc Romeijn , Marcin Chochowski , Mark Cai , Markus Kliegl , Maryam Moosaei , Matt Kulka , Matvei Novikov , Mehrzad Samadi , Melissa Corpuz , Mengru Wang , Meredith Price , Michael Andersch , Michael Boone , Michael Evans , Miguel Martinez , Mikail Khona , Mike Chrzanowski , Minseok Lee , Mohammad Dabbah , Mohammad Shoeybi , Mostofa Patwary , Nabin Mulepati , Najeeb Nabwani , Natalie Hereth , Nave Assaf , Negar Habibi , Neta Zmora , Netanel Haber , Nicola Sessions , Nidhi Bhatia , Nikhil Jukar , Nikki Pope , Nikolai Ludwig , Nima Tajbakhsh , Nir Ailon , Nirmal Juluru , Nishant Sharma , Oleksii Hrinchuk , Oleksii Kuchaiev , Olivier Delalleau , Oluwatobi Olabiyi , Omer Ullman Argov , Omri Puny , Oren Tropp , Ouye Xie , Parth Chadha , Pasha Shamis , Paul Gibbons , Pavlo Molchanov , Pawel Morkisz , Peter Dykas , Peter Jin , Pinky Xu , Piotr Januszewski , Pranav Prashant Thombre , Prasoon Varshney , Pritam Gundecha , Przemek Tredak , Qing Miao , Qiyu Wan , Rabeeh Karimi Mahabadi , Rachit Garg , Ran El-Yaniv , Ran Zilberstein , Rasoul Shafipour , Rich Harang , Rick Izzo , Rima Shahbazyan , Rishabh Garg , Ritika Borkar , Ritu Gala , Riyad Islam , Robert Hesse , Roger Waleffe , Rohit Watve , Roi Koren , Ruoxi Zhang , Russell Hewett , Russell J. Hewett , Ryan Prenger , Ryan Timbrook , Sadegh Mahdavi , Sahil Modi , Samuel Kriman , Sangkug Lim , Sanjay Kariyappa , Sanjeev Satheesh , Saori Kaji , Satish Pasumarthi , Saurav Muralidharan , Sean Narentharen , Sean Narenthiran , Seonmyeong Bak , Sergey Kashirsky , Seth Poulos , Shahar Mor , Shanmugam Ramasamy , Shantanu Acharya , Shaona Ghosh , Sharath Turuvekere Sreenivas , Shelby Thomas , Shiqing Fan , Shreya Gopal , Shrimai Prabhumoye , Shubham Pachori , Shubham Toshniwal , Shuoyang Ding , Siddharth Singh , Simeng Sun , Smita Ithape , Somshubra Majumdar , Soumye Singhal , Stas Sergienko , Stefania Alborghetti , Stephen Ge , Sugam Dipak Devare , Sumeet Kumar Barua , Suseella Panguluri , Suyog Gupta , Sweta Priyadarshi , Syeda Nahida Akter , Tan Bui , Teodor-Dumitru Ene , Terry Kong , Thanh Do , Tijmen Blankevoort , Tim Moon , Tom Balough , Tomer Asida , Tomer Bar Natan , Tomer Ronen , Tugrul Konuk , Twinkle Vashishth , Udi Karpas , Ushnish De , Vahid Noorozi , Vahid Noroozi , Venkat Srinivasan , Venmugil Elango , Victor Cui , Vijay Korthikanti , Vinay Rao , Vitaly Kurin , Vitaly Lavrukhin , Vladimir Anisimov , Wanli Jiang , Wasi Uddin Ahmad , Wei Du , Wei Ping , Wenfei Zhou , Will Jennings , William Zhang , Wojciech Prazuch , Xiaowei Ren , Yashaswi Karnati , Yejin Choi , Yev Meyer , Yi-Fu Wu , Yian Zhang , Yigong Qin , Ying Lin , Yonatan Geifman , Yonggan Fu , Yoshi Subara , Yoshi Suhara , Yubo Gao , Zach Moshe , Zhen Dong , Zhongbo Zhu , Zihan Liu , Zijia Chen , Zijie Yan

Diffusion Large Language Models (dLLMs) offer fast, parallel token generation, but their standalone use is plagued by an inherent efficiency-quality tradeoff. We show that, if carefully applied, the attributes of dLLMs can actually be a…

Machine Learning · Computer Science 2026-01-29 Rui Pan , Zhuofu Chen , Hongyi Liu , Arvind Krishnamurthy , Ravi Netravali

Speculative decoding can significantly accelerate LLM inference, especially given that its cloud-edge collaborative deployment offers cloud workload offloading, offline robustness, and privacy enhancement. However, existing collaborative…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-26 Yunhe Han , Yunqi Gao , Bing Hu , Mahdi Boloursaz Mashhadi , Yitong Duan , Pei Xiao , Yanfeng Zhang

Large Language Models (LLMs) have revolutionized natural language processing by understanding and generating human-like text. However, the increasing demand for more sophisticated LLMs presents significant computational challenges due to…

Computation and Language · Computer Science 2025-01-14 Ze Yang , Yihong Jin , Xinhe Xu

To reduce the latency associated with autoretrogressive LLM inference, speculative decoding has emerged as a novel decoding paradigm, where future tokens are drafted and verified in parallel. However, the practical deployment of speculative…

Computation and Language · Computer Science 2024-12-03 Shwetha Somasundaram , Anirudh Phukan , Apoorv Saxena

Speculative decoding (SD) is a widely adopted approach for accelerating inference in large language models (LLMs), particularly when the draft and target models are well aligned. However, state-of-the-art SD methods typically rely on…

Computation and Language · Computer Science 2026-02-12 Wei Zhong , Manasa Bharadwaj , Yixiao Wang , Yipeng Ji , Chul Lee

Rollout dominates the training time in large language model (LLM) post-training, where the trained model is used to generate tokens given a batch of prompts. This work, SpecActor, achieves fast rollout with speculative decoding that deploys…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-24 Rongxin Cheng , Kai Zhou , Xingda Wei , Siyuan Liu , Mingcong Han , Mingjing Ai , Yeju Zhou , Baoquan Zhong , Wencong Xiao , Rong Chen , Haibo Chen

Speculative decoding has emerged as a promising approach to accelerate autoregressive inference in large language models (LLMs). Self-draft methods, which leverage the base LLM itself for speculation, avoid the overhead of auxiliary draft…

Computation and Language · Computer Science 2026-04-15 Zhuofan Wen , Yang Feng

Inference with modern Large Language Models (LLMs) is expensive and time-consuming, and speculative sampling has proven to be an effective solution. Most speculative sampling methods such as EAGLE use a static draft tree, implicitly…

Computation and Language · Computer Science 2024-07-02 Yuhui Li , Fangyun Wei , Chao Zhang , Hongyang Zhang

Speculative decoding accelerates LLM inference by utilizing otherwise idle computational resources during memory-to-chip data transfer. Current speculative decoding methods typically assume a considerable amount of available computing…

Computation and Language · Computer Science 2025-11-26 Luohe Shi , Zuchao Li , Lefei Zhang , Baoyuan Qi , Guoming Liu , Hai Zhao

Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel with the target model. However, its practical speedup is constrained by the trade-off between draft quality and drafting cost:…

Computation and Language · Computer Science 2026-05-29 Jianuo Huang , Yaojie Zhang , Qituan Zhang , Hao Lin , Hanlin Xu , Linfeng Zhang

Reasoning LLMs produce longer outputs, requiring speculative decoding drafters trained on extended sequences. Parallel drafting - predicting multiple tokens per forward pass - offers latency benefits over sequential generation, but training…

Machine Learning · Computer Science 2026-02-03 Mude Hui , Xin Huang , Jaime Campos Salas , Yue Sun , Nathan Pemberton , Xiang Song , Ashish Khetan , George Karypis

To mitigate the Memory Wall bottleneck encountered by Large Language Models (LLMs) during inference on \textbf{NPU} hardware, and addressing the scarcity of native support for mainstream speculative decoding algorithms on domestic…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-05 Yuntao Dai , Jing Wu , Hang Gu , Teng Wang

Speculative decoding accelerates LLM inference by having a lightweight draft model propose speculative windows of candidate tokens for parallel verification by a larger target model. In practice, speculative efficiency is often bottlenecked…

Computation and Language · Computer Science 2026-05-18 Jie Jiang , Xing Sun , Ruotian Chen , Jianan Su , Kaixin Shen

Recent advancements and widespread adoption of Large Language Models (LLMs) in both industry and academia have catalyzed significant demand for LLM serving. However, traditional cloud services incur high costs, while on-device inference…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-30 Yida Zhang , Zhiyong Gao , Shuaibing Yue , Jie Li , Rui Wang

We introduce AutoJudge, a method that accelerates large language model (LLM) inference with task-specific lossy speculative decoding. Instead of matching the original model output distribution token-by-token, we identify which of the…

Computation and Language · Computer Science 2025-11-21 Roman Garipov , Fedor Velikonivtsev , Ivan Ermakov , Ruslan Svirschevski , Vage Egiazarian , Max Ryabinin

Speculative decoding (SD) accelerates LLM inference by verifying draft tokens in parallel. However, this method presents a critical trade-off: it improves throughput in low-load, memory-bound systems but degrades performance in high-load,…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-04 Rui Li , Zhaoning Zhang , Libo Zhang , Huaimin Wang , Xiang Fu , Zhiquan Lai

As large language models (LLMs) continue to grow in size, distributed inference has become increasingly important. Model-parallel strategies must now efficiently scale not only across multiple GPUs but also across multiple nodes. In this…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-21 Prajwal Singhania , Siddharth Singh , Lannie Dalton Hough , Akarsh Srivastava , Harshitha Menon , Charles Fredrick Jekel , Abhinav Bhatele

In this paper, we introduce a simple training-free technique to improve the performance of drafter-based speculative decoding (SpD) methods that incorporates language modeling head (LM head) during drafting process. A drafter-based…

Computation and Language · Computer Science 2025-09-30 Raghavv Goel , Sudhanshu Agrawal , Mukul Gagrani , Junyoung Park , Yifan Zao , He Zhang , Tian Liu , Yiping Yang , Xin Yuan , Jiuyan Lu , Chris Lott , Mingu Lee

Vision-Language Models (VLMs) excel at visual reasoning but still struggle with integrating external knowledge. Retrieval-Augmented Generation (RAG) is a promising solution, but current methods remain inefficient and often fail to maintain…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Gen Li , Peiyu Liu
‹ Prev 1 3 4 5 6 7 10 Next ›