中文
相关论文

相关论文: RECTOR: Priority-Aware Rule-Based Reranking for Co…

200 篇论文

Reinforcement learning provides an appealing framework for robotic control due to its ability to learn expressive policies purely through real-world interaction. However, this requires addressing real-world constraints and avoiding…

机器人学 · 计算机科学 2024-05-09 Kyle Stachowicz , Sergey Levine

Autonomous vehicles must navigate safely in complex driving environments. Imitating a single expert trajectory, as in regression-based approaches, usually does not explicitly assess the safety of the predicted trajectory. Selection-based…

机器人学 · 计算机科学 2025-11-25 Wenhao Yao , Zhenxin Li , Shiyi Lan , Zi Wang , Xinglong Sun , Jose M. Alvarez , Zuxuan Wu

Language-model agents increasingly emit uncertainty signals throughout a trajectory, but existing agentic UQ evaluations often conflate ranking usefulness with probabilistic truthfulness. AUROC, AUPRC, risk-coverage, Trajectory ECE, and…

人工智能 · 计算机科学 2026-05-26 Suresh Raghu , Satwik Pandey , Shashwat Pandey

Agent-repair leaderboards reorder under evaluator reconfiguration, and a measurable share of the reordering is produced by methods that consult evaluator-derived signal during internal selection of candidate repairs. We document this…

人工智能 · 计算机科学 2026-05-07 Yuelin Hu , Zhenbo Yu , Zhengxue Cheng , Wei Liu , Li Song

Autonomous driving requires safe planning, but most learning-based planners lack explicit self-correction ability: once an unsafe action is proposed, there is no mechanism to correct it. Thus, we propose CorrectionPlanner, an autoregressive…

机器人学 · 计算机科学 2026-03-18 Yihong Guo , Dongqiangzi Ye , Sijia Chen , Anqi Liu , Xianming Liu

Large language models increasingly rely on either reinforcement learning or multi-agent prompting to improve reasoning, yet these two paradigms remain difficult to combine. Directly applying single-agent reinforcement learning to multi-turn…

人工智能 · 计算机科学 2026-05-28 Chusen Li , Zhou Liu , Shuigeng Zhou , Wentao Zhang

Subjective evaluation of LLM behavior -- empathy, restraint, calibrated emotional tone -- is hard. Human inter-rater agreement on such qualities saturates near rho ~ 0.45, and an LLM-as-judge proxy alone risks circularity: a judge sharing…

计算与语言 · 计算机科学 2026-05-28 Yuming , Huang , Yao Liu , Lei Wang , Junchen Wan

Despite recent advances in reinforcement learning (RL), its application in safety critical domains like autonomous vehicles is still challenging. Although punishing RL agents for risky situations can help to learn safe policies, it may also…

机器人学 · 计算机科学 2021-07-16 Danial Kamran , Tizian Engelgeh , Marvin Busch , Johannes Fischer , Christoph Stiller

Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with human scoring standards remains challenging. We formulate this challenge as a criteria-transfer…

计算与语言 · 计算机科学 2026-05-29 Yihan Hong , Huaiyuan Yao , Bolin Shen , Wanpeng Xu , Hua Wei , Yushun Dong

Reinforcement learning (RL) algorithms can achieve state-of-the-art performance in decision-making and continuous control tasks. However, applying RL algorithms on safety-critical systems still needs to be well justified due to the…

机器人学 · 计算机科学 2022-11-22 Mahmoud Selim , Amr Alanwar , M. Watheq El-Kharashi , Hazem M. Abbas , Karl H. Johansson

Reinforcement learning post-training has substantially improved the reasoning accuracy of vision-language models, yet the resulting policies remain poorly calibrated. Terminal correctness rewards provide no gradient that penalizes confident…

机器学习 · 计算机科学 2026-05-19 Peng Cui , Boyao Yang , Jun Zhu

Selective Classification, wherein models can reject low-confidence predictions, promises reliable translation of machine-learning based classification systems to real-world scenarios such as clinical diagnostics. While current evaluation of…

Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the layer where alignment often operates: routing from concept detection to behavioral policy. We…

机器学习 · 计算机科学 2026-05-04 Gregory N. Frank

Humans are remarkably data-efficient when adapting to new unseen conditions, like driving a new car. In contrast, modern robotic control systems, like neural network policies trained using Reinforcement Learning (RL), are highly specialized…

机器人学 · 计算机科学 2026-04-07 Jonas Eschmann , Dario Albani , Giuseppe Loianno

Autonomous driving involves multiple, often conflicting objectives such as safety, efficiency, and comfort. In reinforcement learning (RL), these objectives are typically combined through weighted summation, which collapses their relative…

机器人学 · 计算机科学 2026-03-24 Ahmed Abouelazm , Jonas Michel , Daniel Bogdoll , Philip Schörner , J. Marius Zöllner

Understanding and adhering to traffic regulations is essential for autonomous vehicles to ensure safety and trustworthiness. However, traffic regulations are complex, context-dependent, and differ between regions, posing a major challenge…

人工智能 · 计算机科学 2025-11-20 Tianhui Cai , Yifan Liu , Zewei Zhou , Haoxuan Ma , Seth Z. Zhao , Zhiwen Wu , Xu Han , Zhiyu Huang , Jiaqi Ma

Clinicians need ranking systems that work in real time and still justify their choices. Motivated by the need for a low-latency, decoder-based reranker, we present OG-Rank, a single-decoder approach that pairs a pooled first-token scoring…

人工智能 · 计算机科学 2025-10-21 Praphul Singh , Corey Barrett , Sumana Srivasta , Irfan Bulu , Sri Gadde , Krishnaram Kenthapadi

Reinforcement Learning (RL) has shown exceptional performance across various applications, enabling autonomous agents to learn optimal policies through interaction with their environments. However, traditional RL frameworks often face…

机器学习 · 计算机科学 2025-09-03 Rui Liu , Anish Gupta , Erfaun Noorani , Pratap Tokekar

Navigating complex urban environments safely is a key to realize fully autonomous systems. Predicting future locations of vulnerable road users, such as pedestrians and cyclists, thus, has received a lot of attention in the recent years.…

机器学习 · 计算机科学 2019-10-16 Tessa van der Heiden , Naveen Shankar Nagaraja , Christian Weiss , Efstratios Gavves

Road crashes remain a leading cause of preventable fatalities. Existing prediction models predominantly produce binary outcomes, which offer limited actionable insights for real-time driver feedback. These approaches often lack continuous…

机器学习 · 计算机科学 2026-04-01 Joyjit Roy , Samaresh Kumar Singh , Sushanta Das
‹ 上一页 1 2 3 10 下一页 ›