中文
相关论文

相关论文: CosmoCore Affective Dream-Replay Reinforcement Lea…

200 篇论文

Large language model (LLM) coding agents increasingly operate over repositories, terminals, tests, and execution traces across long software-engineering episodes. Persistent memory is useful, but static vector stores or generic…

软件工程 · 计算机科学 2026-05-05 Mehmet Iscan

Self-correction is a highly desirable capability of large language models (LLMs), yet it has consistently been found to be largely ineffective in modern LLMs. Current methods for training self-correction typically depend on either multiple…

Reinforcement learning (RL) can align language models with non-differentiable reward signals, such as human preferences. However, a major challenge arises from the sparsity of these reward signals - typically, there is only a single reward…

计算与语言 · 计算机科学 2024-02-20 Meng Cao , Lei Shu , Lei Yu , Yun Zhu , Nevan Wichers , Yinxiao Liu , Lei Meng

Large language models excel at complex tasks by breaking down problems into structured reasoning steps. However, reasoning traces often extend beyond reaching a correct answer, causing wasted computation, reduced readability, and…

计算与语言 · 计算机科学 2025-05-26 Razvan-Gabriel Dumitru , Darius Peteleaza , Vikas Yadav , Liangming Pan

Preference-based reward learning is widely used for shaping agent behavior to match a user's preference, yet its sparse binary feedback makes it especially vulnerable to causal confusion. The learned reward often latches onto spurious…

人工智能 · 计算机科学 2026-03-06 Minjune Hwang , Yigit Korkmaz , Daniel Seita , Erdem Bıyık

Large language models (LLMs) exhibit remarkable generative capabilities but often suffer from hallucinations. Retrieval-augmented generation (RAG) offers an effective solution by incorporating external knowledge, but existing methods still…

计算与语言 · 计算机科学 2024-12-17 Xiaoxi Li , Jiajie Jin , Yujia Zhou , Yongkang Wu , Zhonghua Li , Qi Ye , Zhicheng Dou

Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based feedback from unit tests, but its scalability is fundamentally constrained by the availability and…

机器学习 · 计算机科学 2026-05-19 Xiao Zhu , Xinyu Zhou , Boyu Zhu , Hanxu Hu , Mingzhe Du , Haotian Zhang , Huiming Wang , Zhijiang Guo

Large Language Models (LLMs) that can continually improve beyond their training budgets are able to solve increasingly difficult problems by adapting at test time, a property we refer to as extrapolation. However, standard reinforcement…

机器学习 · 计算机科学 2026-03-24 Ian Wu , Yuxiao Qu , Amrith Setlur , Aviral Kumar

Large reasoning models (LRMs) excel on complex problems but face a critical barrier to efficiency: reinforcement learning (RL) training requires long rollouts for outcome-based rewards, where autoregressive decoding dominates time and…

机器学习 · 计算机科学 2026-02-20 Zeliang Zhang , Xiaodong Liu , Hao Cheng , Hao Sun , Chenliang Xu , Jianfeng Gao

Vision-language models (VLMs) raise growing concerns about privacy, copyright, and bias, motivating machine unlearning to remove sensitive knowledge. However, existing methods primarily fine-tune the language decoder, leading to superficial…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Kaidi Jia , Yujie Lin , Chengyi Yang , Jiayao Ma , Jinsong Su

The promise of fault-tolerant quantum computing is challenged by environmental drift that relentlessly degrades the quality of quantum operations. The contemporary solution, halting the entire quantum computation for recalibration, is…

量子物理 · 物理学 2026-03-10 Volodymyr Sivak , Alexis Morvan , Michael Broughton , Rodrigo G. Cortiñas , Johannes Bausch , Andrew W. Senior , Matthew Neeley , Alec Eickbusch , Noah Shutty , Laleh Aghababaie Beni , James S. Spencer , Francisco J. H Heras , Thomas Edlich , Dmitry Abanin , Amira Abbas , Rajeev Acharya , Georg Aigeldinger , Ross Alcaraz , Sayra Alcaraz , Trond I. Andersen , Markus Ansmann , Frank Arute , Kunal Arya , Walt Askew , Nikita Astrakhantsev , Juan Atalaya , Brian Ballard , Joseph C. Bardin , Hector Bates , Andreas Bengtsson , Majid Bigdeli Karimi , Alexander Bilmes , Simon Bilodeau , Felix Borjans , Alexandre Bourassa , Jenna Bovaird , Dylan Bowers , Leon Brill , Peter Brooks , David A. Browne , Brett Buchea , Bob B. Buckley , Tim Burger , Brian Burkett , Nicholas Bushnell , Jamal Busnaina , Anthony Cabrera , Juan Campero , Hung-Shen Chang , Silas Chen , Ben Chiaro , Liang-Ying Chih , Agnetta Y. Cleland , Bryan Cochrane , Matt Cockrell , Josh Cogan , Roberto Collins , Paul Conner , Harold Cook , William Courtney , Alexander L. Crook , Ben Curtin , Martin Damyanov , Sayan Das , Dripto M. Debroy , Sean Demura , Paul Donohoe , Ilya Drozdov , Andrew Dunsworth , Valerie Ehimhen , Aviv Moshe Elbag , Lior Ella , Mahmoud Elzouka , David Enriquez , Catherine Erickson , Vinicius S. Ferreira , Marcos Flores , Leslie Flores Burgos , Ebrahim Forati , Jeremiah Ford , Austin G. Fowler , Brooks Foxen , Masaya Fukami , Alan Wing Lun Fung , Lenny Fuste , Suhas Ganjam , Gonzalo Garcia , Christopher Garrick , Robert Gasca , Helge Gehring , Robert Geiger , Élie Genois , William Giang , Dar Gilboa , James E. Goeders , Edward C. Gonzales , Raja Gosula , Stijn J. de Graaf , Alejandro Grajales Dau , Dietrich Graumann , Joel Grebel , Alex Greene , Jonathan A. Gross , Jose Guerrero , Loïck Le Guevel , Tan Ha , Steve Habegger , Tanner Hadick , Ali Hadjikhani , Michael C. Hamilton , Matthew P. Harrigan , Sean D. Harrington , Jeanne Hartshorn , Stephen Heslin , Paula Heu , Oscar Higgott , Reno Hiltermann , Hsin-Yuan Huang , Mike Hucka , Christopher Hudspeth , Ashley Huff , William J. Huggins , Evan Jeffrey , Shaun Jevons , Zhang Jiang , Xiaoxuan Jin , Chaitali Joshi , Pavol Juhas , Andreas Kabel , Dvir Kafri , Hui Kang , Kiseo Kang , Amir H. Karamlou , Ryan Kaufman , Kostyantyn Kechedzhi , Tanuj Khattar , Mostafa Khezri , Seon Kim , Can M. Knaut , Bryce Kobrin , Fedor Kostritsa , John Mark Kreikebaum , Ryuho Kudo , Ben Kueffler , Arun Kumar , Vladislav D. Kurilovich , Vitali Kutsko , Nathan Lacroix , David Landhuis , Tiano Lange-Dei , Brandon W. Langley , Pavel Laptev , Kim-Ming Lau , Justin Ledford , Joy Lee , Kenny Lee , Brian J. Lester , Wendy Leung , Lily Li , Wing Yan Li , Ming Li , Alexander T. Lill , William P. Livingston , Matthew T. Lloyd , Aditya Locharla , Laura De Lorenzo , Daniel Lundahl , Aaron Lunt , Sid Madhuk , Aniket Maiti , Ashley Maloney , Salvatore Mandrà , Leigh S. Martin , Orion Martin , Eric Mascot , Paul Masih Das , Dmitri Maslov , Melvin Mathews , Cameron Maxfield , Jarrod R. McClean , Matt McEwen , Seneca Meeks , Kevin C. Miao , Zlatko K. Minev , Reza Molavi , Sebastian Molina , Shirin Montazeri , Charles Neill , Michael Newman , Anthony Nguyen , Murray Nguyen , Chia-Hung Ni , Murphy Yuezhen Niu , Logan Oas , Raymond Orosco , Kristoffer Ottosson , Alice Pagano , Agustin Di Paolo , Sherman Peek , David Peterson , Alex Pizzuto , Elias Portoles , Rebecca Potter , Orion Pritchard , Michael Qian , Chris Quintana , Arpit Ranadive , Matthew J. Reagor , Rachel Resnick , David M. Rhodes , Daniel Riley , Gabrielle Roberts , Roberto Rodriguez , Emma Ropes , Lucia B. De Rose , Eliott Rosenberg , Emma Rosenfeld , Dario Rosenstock , Elizabeth Rossi , Pedram Roushan , David A. Rower , Robert Salazar , Kannan Sankaragomathi , Murat Can Sarihan , Kevin J. Satzinger , Max Schaefer , Sebastian Schroeder , Henry F. Schurkus , Aria Shahingohar , Michael J. Shearn , Aaron Shorter , Vladimir Shvarts , Spencer Small , W. Clarke Smith , David A. Sobel , Barrett Spells , Sofia Springer , George Sterling , Jordan Suchard , Aaron Szasz , Alexander Sztein , Madeline Taylor , Jothi Priyanka Thiruraman , Douglas Thor , Dogan Timucin , Eifu Tomita , Alfredo Torres , M. Mert Torunbalci , Hao Tran , Abeer Vaishnav , Justin Vargas , Sergey Vdovichev , Guifre Vidal , Catherine Vollgraff Heidweiller , Meghan Voorhees , Steven Waltman , Jonathan Waltz , Shannon X. Wang , Brayden Ware , James D. Watson , Yonghua Wei , Travis Weidel , Theodore White , Kristi Wong , Bryan W. K. Woo , Christopher J. Wood , Maddy Woodson , Cheng Xing , Z. Jamie Yao , Ping Yeh , Bicheng Ying , Juhwan Yoo , Noureldin Yosri , Elliot Young , Grayson Young , Adam Zalcman , Ran Zhang , Yaxing Zhang , Ningfeng Zhu , Nicholas Zobrist , Zhenjie Zou , Ryan Babbush , Dave Bacon , Sergio Boixo , Yu Chen , Zijun Chen , Michel Devoret , Monica Hansen , Jeremy Hilton , Cody Jones , Julian Kelly , Alexander N. Korotkov , Erik Lucero , Anthony Megrant , Hartmut Neven , William D. Oliver , Ganesh Ramachandran , Vadim Smelyanskiy , Paul V. Klimov

Despite significant advancements in multimodal reasoning tasks, existing Large Vision-Language Models (LVLMs) are prone to producing visually ungrounded responses when interpreting associated images. In contrast, when humans embark on…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Zexian Yang , Dian Li , Dayan Wu , Gang Liu , Weiping Wang

Code Large Language Models (Code LLMs) have demonstrated outstanding performance in code-related tasks. Several instruction tuning approaches have been proposed to boost the code generation performance of pre-trained Code LLMs. In this…

计算与语言 · 计算机科学 2024-02-15 Yejie Wang , Keqing He , Guanting Dong , Pei Wang , Weihao Zeng , Muxi Diao , Yutao Mou , Mengdi Zhang , Jingang Wang , Xunliang Cai , Weiran Xu

While Large Language Models (LLMs) excel at code generation by learning from vast code corpora, a fundamental semantic gap remains between their training on textual patterns and the goal of functional correctness, which is governed by…

Reinforcement Learning (RL) is crucial for empowering VideoLLMs with complex spatiotemporal reasoning. However, current RL paradigms predominantly rely on random data shuffling or naive curriculum strategies based on scalar difficulty…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Hongbo Jin , Kuanwei Lin , Wenhao Zhang , Yichen Jin , Ge Li

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding and generation tasks. However, these models occasionally generate hallucinatory texts, resulting in descriptions that seem reasonable…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Jiaqi Fan , Jianhua Wu , Hongqing Chu , Quanbo Ge , Bingzhao Gao

Reinforcement learning with verifiable rewards (RLVR) has proven effective in eliciting complex reasoning in large language models (LLMs). However, standard RLVR training often leads to excessively verbose processes (in reasoning tasks) and…

人工智能 · 计算机科学 2025-10-01 Gang Li , Yulei Qin , Xiaoyu Tan , Dingkang Yang , Yuchen Shi , Zihan Xu , Xiang Li , Xing Sun , Ke Li

Although people are impressed by the content generation skills of large language models, the use of LLMs, such as ChatGPT, is limited by the domain grounding of the content. The correctness and groundedness of the generated content need to…

计算与语言 · 计算机科学 2024-12-23 Xiaofeng Zhu , Jaya Krishna Mandivarapu

While multimodal reasoning models (MLRMs) have exhibited impressive capabilities, they remain prone to hallucinations, and effective solutions are still underexplored. In this paper, we experimentally analyze the hallucination cause and…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Hao Fang , Jinyu Li , Jiawei Kong , Tianqu Zhuang , Kuofeng Gao , Bin Chen , Shu-Tao Xia

Large language models (LLMs) can generate executable code from natural language descriptions, but the resulting programs frequently contain bugs due to hallucinations. In the absence of formal specifications, existing approaches attempt to…

软件工程 · 计算机科学 2026-03-31 Yihan Dai , Sijie Liang , Haotian Xu , Peichu Xie , Sergey Mechtaev