中文
相关论文

相关论文: Closed-Loop Supervised Fine-Tuning of Tokenized Tr…

200 篇论文

In this article, we explore the feasibility of applying proximal policy optimization, a state-of-the-art deep reinforcement learning algorithm for continuous control tasks, on the dual-objective problem of controlling an underactuated…

机器学习 · 计算机科学 2019-12-20 Eivind Meyer , Haakon Robinson , Adil Rasheed , Omer San

It is known that reinforcement learning (RL) is data-hungry. To improve sample-efficiency of RL, it has been proposed that the learning algorithm utilize data from 'approximately similar' processes. However, since the process models are…

机器学习 · 计算机科学 2025-11-24 Vinay Kanakeri , Shivam Bajaj , Ashwin Verma , Vijay Gupta , Aritra Mitra

Fine-tuning simulation-trained RL agents with real-world data often degrades crucial behaviors due to limited or skewed data distributions. We argue that designer priorities exist not just in reward functions, but also in simulation design…

机器人学 · 计算机科学 2025-05-02 Bassel El Mabsout , Shahin Roozkhosh , Siddharth Mysore , Kate Saenko , Renato Mancuso

Optimal traffic-light settings are generally hard to obtain, certainly for actuated access control of an intersection. Typically, computationally expensive (microscopic) simulations or complicated optimization schemes are required to find…

信号处理 · 电气工程与系统科学 2020-07-17 Rik W. Timmerman , Marko A. A. Boon

Cooperative maneuver planning promises to significantly improve traffic efficiency at unsignalized intersections by leveraging connected automated vehicles. Previous works on this topic have been mostly developed for completely automated…

机器人学 · 计算机科学 2026-02-03 Marvin Klimke , Max Bastian Mertens , Benjamin Völz , Michael Buchholz

Traffic Signal Control (TSC) is essential for managing urban traffic flow and reducing congestion. Reinforcement Learning (RL) offers an adaptive method for TSC by responding to dynamic traffic patterns, with multi-agent RL (MARL) gaining…

机器学习 · 计算机科学 2025-07-22 Justin Turnau , Longchao Da , Khoa Vo , Ferdous Al Rafi , Shreyas Bachiraju , Tiejin Chen , Hua Wei

Traffic simulation provides interactive data for the optimization of traffic control policies. However, existing traffic simulators are limited by their lack of scalability and shortage in input data, which prevents them from generating…

Modern approaches to autonomous driving rely heavily on learned components trained with large amounts of human driving data via imitation learning. However, these methods require large amounts of expensive data collection and even then face…

In this paper, methods have been explored to effectively optimise traffic signal control to minimise waiting times and queue lengths, thereby increasing traffic flow. The traffic intersection was first defined as a Markov Decision Process,…

系统与控制 · 电气工程与系统科学 2022-07-29 Hrishit Chaudhuri , Vibha Masti , Vishruth Veerendranath , S Natarajan

The recent advancements in wireless technology enable connected autonomous vehicles (CAVs) to gather information about their environment by vehicle-to-vehicle (V2V) communication. In this work, we design an information-sharing-based…

人工智能 · 计算机科学 2022-09-07 Songyang Han , Shanglin Zhou , Jiangwei Wang , Lynn Pepin , Caiwen Ding , Jie Fu , Fei Miao

Behavior Cloning (BC) on curated (or filtered) data is the predominant paradigm for supervised fine-tuning (SFT) of large language models; as well as for imitation learning of control policies. Here, we draw on a connection between this…

机器学习 · 计算机科学 2025-09-09 Chongli Qin , Jost Tobias Springenberg

While the capabilities of autonomous driving have advanced rapidly, merging into dense traffic remains a significant challenge, many motion planning methods for this scenario have been proposed but it is hard to evaluate them. Most existing…

机器人学 · 计算机科学 2025-04-03 Zhengming Wang , Junli Wang , Pengfei Li , Zhaohan Li , Chunyang Liu , Bo Zhang , Peng Li , Yilun Chen

Collaborative navigation becomes essential in situations of occluded scenarios in autonomous driving where independent driving policies are likely to lead to collisions. One promising approach to address this issue is through the use of…

机器人学 · 计算机科学 2024-12-12 Leandro Parada , Hanlin Tian , Jose Escribano , Panagiotis Angeloudis

The number of large language models (LLMs) with varying parameter scales and vocabularies is increasing. While they deliver powerful performance, they also face a set of common optimization needs to meet specific requirements or standards,…

计算与语言 · 计算机科学 2024-10-24 Jiayi Wu , Hao Sun , Hengyi Cai , Lixin Su , Shuaiqiang Wang , Dawei Yin , Xiang Li , Ming Gao

Recently, safe reinforcement learning (RL) with the actor-critic structure for continuous control tasks has received increasing attention. It is still challenging to learn a near-optimal control policy with safety and convergence…

机器学习 · 计算机科学 2024-02-06 Xinglong Zhang , Yaoqian Peng , Biao Luo , Wei Pan , Xin Xu , Haibin Xie

Reinforcement learning (RL) has shown promise in robotics, but deploying RL on real vehicles remains challenging due to the complexity of vehicle dynamics and the mismatch between simulation and reality. Factors such as tire…

机器人学 · 计算机科学 2025-11-11 Thomas Steinecker , Alexander Bienemann , Denis Trescher , Thorsten Luettel , Mirko Maehlisch

Existing GUI agent models relying on coordinate-based one-step visual grounding struggle with generalizing to varying input resolutions and aspect ratios. Alternatives introduce coordinate-free strategies yet suffer from learning under…

机器学习 · 计算机科学 2026-02-04 Xiaoce Wang , Guibin Zhang , Junzhe Li , Jinzhe Tu , Chun Li , Ming Li

This paper investigates how the performance of visual navigation policies trained in simulation compares to policies trained with real-world data. Performance degradation of simulator-trained policies is often significant when they are…

Imitation learning is a control design paradigm that seeks to learn a control policy reproducing demonstrations from expert agents. By substituting expert demonstrations for optimal behaviours, the same paradigm leads to the design of…

机器学习 · 计算机科学 2024-12-20 Dharmesh Tailor , Dario Izzo

We present simulations of congested traffic in circular and open systems with a non-local, gas-kinetic-based traffic model and a novel car-following model. The model parameters are all intuitive and can be easily calibrated. Micro- and…

统计力学 · 物理学 2007-05-23 Dirk Helbing , Ansgar Hennecke , Vladimir Shvetsov , Martin Treiber
‹ 上一页 1 8 9 10 下一页 ›