中文
相关论文

相关论文: A/B Testing Measurement Framework for Recommendati…

200 篇论文

We consider the problem of designing prices for public transport where payment enforcing is done through random inspection of passengers' tickets as opposed to physically blocking their access. Passengers are fully strategic such that they…

理论经济学 · 经济学 2026-04-27 Inácio Bó , Chiu Yu Ko

Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning abilities of Large Language Models (LLMs) by using rule-based binary feedback. However, current RLVR methods typically assign the same reward to every token.…

机器学习 · 计算机科学 2025-10-21 Guofu Xie , Yunsheng Shi , Hongtao Tian , Ting Yao , Xiao Zhang

The evaluation of recommender systems from a practical perspective is a topic of ongoing discourse within the research community. While many current evaluation methods reduce performance to a single value metric as an easy way to compare…

机器学习 · 计算机科学 2023-02-13 Mikhail Andronov , Sergey Kolesnikov

In this work, we study an upgrading scheme for online resource allocation problems. We work in a sequential setting, where at each round a request for a resource arrives and the decision-maker has to decide whether to accept it (and thus,…

最优化与控制 · 数学 2024-02-15 Patrick Jaillet , Chara Podimata , Andrew Vakhutinsky , Zijie Zhou

Offline-to-Online Reinforcement Learning (O2O RL) faces a critical dilemma in balancing the use of a fixed offline dataset with newly collected online experiences. Standard methods, often relying on a fixed data-mixing ratio, struggle to…

机器学习 · 计算机科学 2026-04-09 Chihyeon Song , Jaewoo Lee , Jinkyoo Park

The Linear Parameter-Varying (LPV) framework has long been used to guarantee performance and stability requirements of nonlinear (NL) systems mainly through the $\mathcal{L}_2$-gain concept. However, recent research has pointed out that…

系统与控制 · 电气工程与系统科学 2020-05-14 P. J. W. Koelewijn , R. Tóth , H. Nijmeijer

Over the past decade, most technology companies and a growing number of conventional firms have adopted online experimentation (or A/B testing) into their product development process. Initially, A/B testing was deployed as a static…

应用统计 · 统计学 2021-11-04 Jialiang Mao , Iavor Bojinov

Though it has been recognized that recommending serendipitous (i.e., surprising and relevant) items can be helpful for increasing users' satisfaction and behavioral intention, how to measure serendipity in the offline environment is still…

人机交互 · 计算机科学 2020-04-23 Li Chen , Ningxia Wang , Yonghua Yang , Keping Yang , Quan Yuan

Real-time bidding (RTB) has become one of the largest online advertising markets in the world. Today the bid price per ad impression is typically decided by the expected value of how it can lead to a desired action event (e.g., registering…

计算机科学与博弈论 · 计算机科学 2016-02-16 Jian Xu , Xuhui Shao , Jianjie Ma , Kuang-chih Lee , Hang Qi , Quan Lu

In this paper, we examine the biases that arise when firms run A/B tests on continuous parameters to estimate global treatment effects on performance metrics of interest; we particularly focus on price experiments to measure the price…

统计方法学 · 统计学 2026-01-22 Ramesh Johari , Orrie B. Page , Gabriel Y. Weintraub

We propose a novel algorithm for offline reinforcement learning called Value Iteration with Perturbed Rewards (VIPeR), which amalgamates the pessimism principle with random perturbations of the value function. Most current offline RL…

机器学习 · 计算机科学 2023-03-07 Thanh Nguyen-Tang , Raman Arora

The paper proposes and optimizes a partial recovery training system, CPR, for recommendation models. CPR relaxes the consistency requirement by enabling non-failed nodes to proceed without loading checkpoints when a node fails during…

Modern recommender systems aim to improve user experience. As reinforcement learning (RL) naturally fits this objective -- maximizing an user's reward per session -- it has become an emerging topic in recommender systems. Developing…

Recommender systems can be formulated as a matrix completion problem, predicting ratings from user and item parameter vectors. Optimizing these parameters by subsampling data becomes difficult as the number of users and items grows. We…

信息检索 · 计算机科学 2018-07-09 Elias Tragas , Calvin Luo , Maxime Gazeau , Kevin Luk , David Duvenaud

Existing methods for extracting reward signals in Reinforcement Learning typically rely on labeled data and dedicated training splits, a setup that contrasts with how humans learn directly from their environment. In this work, we propose…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Akshit Singh , Shyam Marjit , Wei Lin , Paul Gavrikov , Serena Yeung-Levy , Hilde Kuehne , Rogerio Feris , Sivan Doveh , James Glass , M. Jehanzeb Mirza

We introduce AVA, an automatic evaluation approach for Question Answering, which given a set of questions associated with Gold Standard answers, can estimate system Accuracy. AVA uses Transformer-based language models to encode question,…

计算与语言 · 计算机科学 2021-08-17 Thuy Vu , Alessandro Moschitti

It is increasingly common in digital environments to use A/B tests to compare the performance of recommendation algorithms. However, such experiments often violate the stable unit treatment value assumption (SUTVA), particularly SUTVA's "no…

Rewards serve as a measure of user satisfaction and act as a limiting factor in interactive recommender systems. In this research, we focus on the problem of learning to reward (LTR), which is fundamental to reinforcement learning. Previous…

机器学习 · 计算机科学 2023-10-31 Jialin Liu , Xinyan Su , Zeyu He , Xiangyu Zhao , Jun Li

Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective post-training paradigm for improving the reasoning capabilities of large language models. However, existing group-based RLVR methods often suffer from severe…

机器学习 · 计算机科学 2026-05-26 Haechan Kim , Soohyun Ryu , Gyouk Chu , Doohyuk Jang , Eunho Yang

Multimodal recommendation faces an issue of the performance degradation that the uni-modal recommendation sometimes achieves the better performance. A possible reason is that the unreliable item modality data hurts the fusion result.…

信息检索 · 计算机科学 2025-04-24 Xue Dong , Xuemeng Song , Na Zheng , Sicheng Zhao , Guiguang Ding
‹ 上一页 1 8 9 10 下一页 ›