中文
相关论文

相关论文: HALO: Hindsight-Augmented Learning for Online Auto…

200 篇论文

Human-in-the-loop Bayesian optimization (HITL BO) methods utilize human expertise to improve the sample-efficiency of BO. Most HITL BO methods assume that a domain expert can quantify their knowledge, for instance by pinpointing query…

机器学习 · 计算机科学 2026-05-13 Alvar Haltia , Ville Hyvönen , Samuel Kaski

In this work, we propose a hierarchical reinforcement learning (HRL) structure which is capable of performing autonomous vehicle planning tasks in simulated environments with multiple sub-goals. In this hierarchical structure, the network…

机器人学 · 计算机科学 2019-11-12 Zhiqian Qiao , Zachariah Tyree , Priyantha Mudalige , Jeff Schneider , John M. Dolan

Real-time Bidding (RTB) advertisers wish to \textit{know in advance} the expected cost and yield of ad campaigns to avoid trial-and-error expenses. However, Campaign Performance Forecasting (CPF), a sequence modeling task involving tens of…

信息检索 · 计算机科学 2024-05-20 XiaoYu Wang , YongHui Guo , Hui Sheng , Peili Lv , Chi Zhou , Wei Huang , ShiQin Ta , Dongbo Huang , XiuJin Yang , Lan Xu , Hao Zhou , Yusheng Ji

Hyperparameter optimization (HPO) is a billion-dollar problem in machine learning, which significantly impacts the training efficiency and model performance. However, achieving efficient and robust HPO in deep reinforcement learning (RL) is…

机器学习 · 计算机科学 2025-08-04 Mingqi Yuan , Bo Li , Xin Jin , Wenjun Zeng

Today's online advertisers procure digital ad impressions through interacting with autobidding platforms: advertisers convey high level procurement goals via setting levers such as budget, target return-on-investment, max cost per click,…

信息检索 · 计算机科学 2023-07-13 Jason Cheuk Nam Liang , Haihao Lu , Baoyu Zhou

Meta-reinforcement learning (meta-RL) algorithms allow for agents to learn new behaviors from small amounts of experience, mitigating the sample inefficiency problem in RL. However, while meta-RL agents can adapt quickly to new tasks at…

机器学习 · 计算机科学 2022-04-26 Michael Wan , Jian Peng , Tanmay Gangwani

Online auction scenarios, such as bidding searches on advertising platforms, often require bidders to participate repeatedly in auctions for identical or similar items. Most previous studies have only considered the process by which the…

计算机科学与博弈论 · 计算机科学 2024-02-28 Yudong Hu , Congying Han , Tiande Guo , Hao Xiao

Humanoid robots deployed in real-world scenarios often need to carry unknown payloads, which introduce significant mismatch and degrade the effectiveness of simulation-to-reality reinforcement learning methods. To address this challenge, we…

机器人学 · 计算机科学 2026-03-17 Xingyi Wang , Chenyun Zhang , Weiji Xie , Chao Yu , Wei Song , Chenjia Bai , Shiqiang Zhu

Modern content platforms offer paid promotion to mitigate cold start by allocating exposure via auctions. Our empirical analysis reveals a counterintuitive flaw in this paradigm: while promotion rescues low-to-medium quality content, it can…

计算机科学与博弈论 · 计算机科学 2026-01-29 Yumou Liu , Zhenzhe Zheng , Jiang Rong , Yao Hu , Fan Wu , Guihai Chen

With the emergence of new online channels and information technology, digital advertising tends to substitute more and more to traditional advertising by offering the opportunity to companies to target the consumers/users that are really…

最优化与控制 · 数学 2021-11-17 Médéric Motte , Huyên Pham

The emergence of real-time auction in online advertising has drawn huge attention of modeling the market competition, i.e., bid landscape forecasting. The problem is formulated as to forecast the probability distribution of market price for…

信息检索 · 计算机科学 2019-05-14 Kan Ren , Jiarui Qin , Lei Zheng , Zhengyu Yang , Weinan Zhang , Yong Yu

We introduce a novel large-scale deep learning model for Limit Order Book mid-price changes forecasting, and we name it `HLOB'. This architecture (i) exploits the information encoded by an Information Filtering Network, namely the…

交易与市场微观结构 · 定量金融 2024-06-05 Antonio Briola , Silvia Bartolucci , Tomaso Aste

An important challenge in metric learning is scalability to both size and dimension of input data. Online metric learning algorithms are proposed to address this challenge. Existing methods are commonly based on (Passive Aggressive) PA…

机器学习 · 计算机科学 2020-10-13 Davood Zabihzadeh , Amar Tuama , Ali Karami-Mollaee

Real-Time Bidding (RTB) is revolutionising display advertising by facilitating per-impression auctions to buy ad impressions as they are being generated. Being able to use impression-level data, such as user cookies, encourages user…

计算机科学与博弈论 · 计算机科学 2016-03-04 Weinan Zhang , Yifei Rong , Jun Wang , Tianchi Zhu , Xiaofan Wang

Ad exchanges are widely used in platforms for online display advertising. Autonomous agents operating in these exchanges must learn policies for interacting profitably with a diverse, continually changing, but unknown market. We consider…

计算机科学与博弈论 · 计算机科学 2019-02-12 Stavros Gerakaris , Subramanian Ramamoorthy

We study the problem of selecting large language models (LLMs) for user queries in settings where multiple LLM providers submit the cost of solving a query. From the users' perspective, choosing an optimal model is a sequential,…

计算机科学与博弈论 · 计算机科学 2026-02-17 Pronoy Patra , Sankarshan Damle , Manisha Padala , Sujit Gujar

High-frequency trading (HFT) that executes algorithmic trading in short time scales, has recently occupied the majority of cryptocurrency market. Besides traditional quantitative trading methods, reinforcement learning (RL) has become…

机器学习 · 计算机科学 2024-06-21 Chuqiao Zong , Chaojie Wang , Molei Qin , Lei Feng , Xinrun Wang , Bo An

Hierarchical Instruction Following (HIF) refers to the problem of prompting large language models with a priority-ordered stack of instructions. Standard methods like RLHF and DPO typically fail in this problem since they mainly optimize…

机器学习 · 计算机科学 2026-03-18 Keru Chen , Jun Luo , Sen Lin , Yingbin Liang , Alvaro Velasquez , Nathaniel Bastian , Shaofeng Zou

Auctions are important mechanisms extensively implemented in various markets, e.g., search engines' keyword auctions, antique auctions, etc. Finding an optimal auction mechanism is extremely difficult due to the constraints of imperfect…

机器学习 · 计算机科学 2025-07-28 Jiayin Liu , Chenglong Zhang

Hyperparameter optimization (HPO) plays a central role in the performance of deep learning models, yet remains computationally expensive and difficult to interpret, particularly for time-series forecasting. While Bayesian Optimization (BO)…

机器学习 · 计算机科学 2026-02-17 Ons Saadallah , Mátyás andó , Tamás Gábor Orosz