中文
相关论文

相关论文: UFO-RL: Uncertainty-Focused Optimization for Effic…

200 篇论文

Due to their adaptability and mobility, Unmanned Aerial Vehicles (UAVs) are becoming increasingly essential for wireless network services, particularly for data harvesting tasks. In this context, Artificial Intelligence (AI)-based…

机器学习 · 计算机科学 2026-01-21 Babacar Toure , Dimitrios Tsilimantos , Omid Esrafilian , Marios Kountouris

Deployment efficiency is an important criterion for many real-world applications of reinforcement learning (RL). Despite the community's increasing interest, there lacks a formal theoretical formulation for the problem. In this paper, we…

机器学习 · 计算机科学 2022-09-01 Jiawei Huang , Jinglin Chen , Li Zhao , Tao Qin , Nan Jiang , Tie-Yan Liu

As models grow larger and training them becomes expensive, it becomes increasingly important to scale training recipes not just to larger models and more data, but to do so in a compute-optimal manner that extracts maximal performance per…

机器学习 · 计算机科学 2025-08-26 Preston Fu , Oleh Rybkin , Zhiyuan Zhou , Michal Nauman , Pieter Abbeel , Sergey Levine , Aviral Kumar

Reinforcement learning (RL) has emerged as a promising strategy for improving the reasoning capabilities of language models (LMs) in domains such as mathematics and coding. However, most modern RL algorithms were designed to target robotics…

人工智能 · 计算机科学 2025-05-26 Lianghuan Huang , Shuo Li , Sagnik Anupam , Insup Lee , Osbert Bastani

Safe reinforcement learning (Safe RL) refers to a class of techniques that aim to prevent RL algorithms from violating constraints in the process of decision-making and exploration during trial and error. In this paper, a novel model-free…

系统与控制 · 电气工程与系统科学 2024-08-14 Homayoun Honari , Mehran Ghafarian Tamizi , Homayoun Najjaran

Modern deep architectures often rely on large-scale datasets, but training on these datasets incurs high computational and storage overhead. Real-world datasets often contain substantial redundancies, prompting the need for more…

机器学习 · 计算机科学 2025-06-27 Suorong Yang , Peijia Li , Furao Shen , Jian Zhao

Cross-domain offline reinforcement learning (RL) aims to train a well-performing agent in the target environment, leveraging both a limited target domain dataset and a source domain dataset with (possibly) sufficient data coverage. Due to…

机器学习 · 计算机科学 2026-03-23 Zhongjian Qiao , Rui Yang , Jiafei Lyu , Chenjia Bai , Xiu Li , Siyang Gao , Shuang Qiu

Over recent years, an increasing amount of compute and data has been poured into training large language models (LLMs), usually by doing one-pass learning on as many tokens as possible randomly selected from large-scale web corpora. While…

计算与语言 · 计算机科学 2023-08-24 Kushal Tirumala , Daniel Simig , Armen Aghajanyan , Ari S. Morcos

Contemporary autopilot systems for unmanned aerial vehicles (UAVs) are far more limited in their flight envelope as compared to experienced human pilots, thereby restricting the conditions UAVs can operate in and the types of missions they…

机器人学 · 计算机科学 2019-11-14 Eivind Bøhn , Erlend M. Coates , Signe Moe , Tor Arne Johansen

Improving sampling efficiency and generalization capability is critical for the successful data-driven control of quadrotor unmanned aerial vehicles (UAVs) that are inherently unstable. While various reinforcement learning (RL) approaches…

机器人学 · 计算机科学 2025-03-03 Beomyeol Yu , Taeyoung Lee

The learning inefficiency of reinforcement learning (RL) from scratch hinders its practical application towards continuous robotic tracking control, especially for high-dimensional robots. This work proposes a data-informed residual…

系统与控制 · 电气工程与系统科学 2024-06-10 Cong Li , Fangzhou Liu , Yongchao Wang , Martin Buss

In robot manipulation, Reinforcement Learning (RL) often suffers from low sample efficiency and uncertain convergence, especially in large observation and action spaces. Foundation Models (FMs) offer an alternative, demonstrating promise in…

机器人学 · 计算机科学 2025-04-18 Runyu Ma , Jelle Luijkx , Zlatan Ajanovic , Jens Kober

Reinforcement learning (RL) algorithms aim to learn optimal decisions in unknown environments through experience of taking actions and observing the rewards gained. In some cases, the environment is not influenced by the actions of the RL…

Retrieval-Augmented Generation (RAG) extends Large Vision-Language Models (LVLMs) with external visual knowledge. However, existing visual RAG systems typically rely on generic retrieval signals that overlook the fine-grained visual…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Jun Wang , Shuo Tan , Zelong Sun , Tiancheng Gu , Yongle Zhao , Ziyong Feng , Kaicheng Yang , Zhiwu Lu

While reinforcement learning (RL) holds great potential for decision making in the real world, it suffers from a number of unique difficulties which often need specific consideration. In particular: it is highly non-stationary; suffers from…

Reinforcement Learning (RL) heavily relies on the careful design of the reward function. However, accurately assigning rewards to each state-action pair in Long-Term Reinforcement Learning (LTRL) tasks remains a significant challenge. As a…

机器学习 · 计算机科学 2025-06-03 Qi Ju , Falin Hei , Zhemei Fang , Yunfeng Luo

Large language models are increasingly used for complex reasoning tasks where high-quality offline data such as expert-annotated solutions and distilled reasoning traces are often available. However, in environments with sparse rewards,…

人工智能 · 计算机科学 2025-08-11 Yihao Liu , Shuocheng Li , Lang Cao , Yuhang Xie , Mengyu Zhou , Haoyu Dong , Xiaojun Ma , Shi Han , Dongmei Zhang

Safe flight in dynamic environments requires unmanned aerial vehicles (UAVs) to make effective decisions when navigating cluttered spaces with moving obstacles. Traditional approaches often decompose decision-making into hierarchical…

机器人学 · 计算机科学 2025-02-25 Zhefan Xu , Xinming Han , Haoyu Shen , Hanyu Jin , Kenji Shimada

In the era of large-scale surveys like Euclid, machine learning has become an essential tool for identifying rare yet scientifically valuable objects, such as strong gravitational lenses. However, supervised machine-learning approaches…

天体物理仪器与方法 · 物理学 2025-12-08 Euclid Collaboration , N. E. P. Lines , T. E. Collett , P. Holloway , K. Rojas , S. Schuldt , R. B. Metcalf , T. Li , A. Verma , G. Despali , F. Courbin , R. Gavazzi , C. Tortora , B. Clément , N. Aghanim , B. Altieri , L. Amendola , S. Andreon , N. Auricchio , C. Baccigalupi , M. Baldi , A. Balestra , S. Bardelli , P. Battaglia , A. Biviano , E. Branchini , M. Brescia , S. Camera , G. Cañas-Herrera , V. Capobianco , C. Carbone , J. Carretero , M. Castellano , G. Castignani , S. Cavuoti , A. Cimatti , C. Colodro-Conde , G. Congedo , C. J. Conselice , L. Conversi , Y. Copin , H. M. Courtois , M. Cropper , H. Degaudenzi , G. De Lucia , H. Dole , F. Dubath , X. Dupac , S. Dusini , A. Ealet , S. Escoffier , M. Farina , R. Farinelli , F. Faustini , S. Ferriol , F. Finelli , M. Frailis , E. Franceschi , M. Fumana , S. Galeotta , K. George , B. Gillis , C. Giocoli , P. Gómez-Alvarez , J. Gracia-Carpio , A. Grazian , F. Grupp , S. V. H. Haugan , W. Holmes , I. M. Hook , F. Hormuth , A. Hornstrup , K. Jahnke , M. Jhabvala , B. Joachimi , E. Keihänen , S. Kermiche , A. Kiessling , B. Kubik , M. Kümmel , M. Kunz , H. Kurki-Suonio , A. M. C. Le Brun , S. Ligori , P. B. Lilje , V. Lindholm , I. Lloro , G. Mainetti , D. Maino , E. Maiorano , O. Mansutti , S. Marcin , O. Marggraf , M. Martinelli , N. Martinet , F. Marulli , R. J. Massey , E. Medinaceli , S. Mei , M. Melchior , Y. Mellier , M. Meneghetti , E. Merlin , G. Meylan , A. Mora , M. Moresco , L. Moscardini , R. Nakajima , C. Neissner , S. -M. Niemi , J. W. Nightingale , C. Padilla , S. Paltani , F. Pasian , K. Pedersen , W. J. Percival , V. Pettorino , S. Pires , G. Polenta , M. Poncet , L. A. Popa , L. Pozzetti , F. Raison , A. Renzi , J. Rhodes , G. Riccio , E. Romelli , M. Roncarelli , C. Rosset , R. Saglia , Z. Sakr , A. G. Sánchez , D. Sapone , B. Sartoris , J. A. Schewtschenko , P. Schneider , T. Schrabback , A. Secroun , G. Seidel , S. Serrano , C. Sirignano , G. Sirri , L. Stanco , J. Steinwagner , P. Tallada-Crespí , A. N. Taylor , I. Tereno , N. Tessore , S. Toft , R. Toledo-Moreo , F. Torradeflot , I. Tutusaus , J. Valiviita , T. Vassallo , A. Veropalumbo , Y. Wang , J. Weller , A. Zacchei , G. Zamorani , F. M. Zerbi , E. Zucca , M. Ballardini , M. Bolzonella , E. Bozzo , C. Burigana , R. Cabanac , M. Calabrese , A. Cappi , T. Castro , J. A. Escartin Vigo , L. Gabarra , J. García-Bellido , V. Gautard , S. Hemmati , M. Huertas-Company , J. Macias-Perez , R. Maoli , J. Martín-Fleitas , M. Maturi , N. Mauri , P. Monaco , M. Pöntinen , C. Porciani , I. Risso , V. Scottez , M. Sereno , M. Tenti , M. Tucci , M. Viel , M. Wiesmann , Y. Akrami , I. T. Andika , G. Angora , S. Anselmi , M. Archidiacono , F. Atrio-Barandela , E. Aubourg , L. Bazzanini , D. Bertacca , M. Bethermin , F. Beutler , A. Blanchard , L. Blot , M. Bonici , S. Borgani , M. L. Brown , S. Bruton , A. Calabro , B. Camacho Quevedo , F. Caro , C. S. Carvalho , F. Cogato , S. Conseil , A. R. Cooray , O. Cucciati , S. Davini , F. De Paolis , G. Desprez , A. Díaz-Sánchez , S. Di Domizio , J. M. Diego , P. -A. Duc , V. Duret , M. Y. Elkhashab , A. Enia , Y. Fang , P. G. Ferreira , A. Finoguenov , A. Fontana , A. Franco , K. Ganga , T. Gasparetto , E. Gaztanaga , F. Giacomini , F. Gianotti , G. Gozaliasl , A. Gruppuso , M. Guidi , C. M. Gutierrez , A. Hall , H. Hildebrandt , J. Hjorth , J. J. E. Kajava , Y. Kang , V. Kansal , D. Karagiannis , K. Kiiveri , J. Kim , C. C. Kirkpatrick , S. Kruk , M. Lattanzi , L. Legrand , F. Lepori , G. Leroy , G. F. Lesci , J. Lesgourgues , T. I. Liaudat , M. Magliocchetti , A. Manjón-García , F. Mannucci , C. J. A. P. Martins , L. Maurin , M. Miluzio , A. Montoro , C. Moretti , G. Morgante , S. Nadathur , K. Naidoo , P. Natoli , S. Nesseris , D. Paoletti , F. Passalacqua , K. Paterson , L. Patrizii , A. Pisani , D. Potter , G. W. Pratt , S. Quai , M. Radovich , W. Roster , S. Sacquegna , M. Sahlén , D. B. Sanders , E. Sarpa , A. Schneider , D. Sciotti , E. Sellentin , L. C. Smith , J. G. Sorce , K. Tanidis , C. Tao , F. Tarsitano , G. Testera , R. Teyssier , S. Tosi , A. Troja , A. Venhola , D. Vergani , G. Vernardos , G. Verza , S. Vinciguerra , M. Walmsley , N. A. Walton , A. H. Wright

Fine-tuning Large Language Models (LLMs) typically relies on large quantities of high-quality annotated data, or questions with well-defined ground truth answers in the case of Reinforcement Learning with Verifiable Rewards (RLVR). While…

人工智能 · 计算机科学 2026-04-21 Justin Bauer , Thomas Walshe , Derek Pham , Harit Vishwakarma , Armin Parchami , Frederic Sala , Paroma Varma