English
Related papers

Related papers: CASP: Support-Aware Offline Policy Selection for T…

200 papers

We study online control for continuous-time linear systems with finite sampling rates, where the objective is to design an online procedure that learns under non-stochastic noise and performs comparably to a fixed optimal linear controller.…

Optimization and Control · Mathematics 2025-06-10 Jingwei Li , Jing Dong , Can Chang , Baoxiang Wang , Jingzhao Zhang

In offline RL, constraining the learned policy to remain close to the data is essential to prevent the policy from outputting out-of-distribution (OOD) actions with erroneously overestimated values. In principle, generative adversarial…

Machine Learning · Computer Science 2022-11-04 Quan Vuong , Aviral Kumar , Sergey Levine , Yevgen Chebotar

We primarily focus on the field of multi-scenario recommendation, which poses a significant challenge in effectively leveraging data from different scenarios to enhance predictions in scenarios with limited data. Current mainstream efforts…

Information Retrieval · Computer Science 2024-04-16 Jiachen Zhu , Yichao Wang , Jianghao Lin , Jiarui Qin , Ruiming Tang , Weinan Zhang , Yong Yu

Model-based algorithms, which learn a dynamics model from logged experience and perform some sort of pessimistic planning under the learned model, have emerged as a promising paradigm for offline reinforcement learning (offline RL).…

Machine Learning · Computer Science 2022-01-28 Tianhe Yu , Aviral Kumar , Rafael Rafailov , Aravind Rajeswaran , Sergey Levine , Chelsea Finn

Retrieval-Augmented Generation (RAG) improves generation quality by incorporating evidence retrieved from large external corpora. However, most existing methods rely on statically selecting top-k passages based on individual relevance,…

Artificial Intelligence · Computer Science 2026-01-09 Yi Jiang , Sendong Zhao , Jianbo Li , Bairui Hu , Yanrui Du , Haochun Wang , Bing Qin

In practical applications of semantic parsing, we often want to rapidly change the behavior of the parser, such as enabling it to handle queries in a new domain, or changing its predictions on certain targeted queries. While we can…

Computation and Language · Computer Science 2022-02-24 Panupong Pasupat , Yuan Zhang , Kelvin Guu

Large-scale recommender systems often face severe latency and storage constraints at prediction time. These are particularly acute when the number of items that could be recommended is large, and calculating predictions for the full set is…

Information Retrieval · Computer Science 2017-09-05 Maciej Kula

For a data-generating process for random variables that can be described with a linear structural equation model, we consider a situation in which (i) a set of covariates satisfying the back-door criterion cannot be observed or (ii) such a…

Methodology · Statistics 2025-03-06 Hisayoshi Nanmo , Manabu Kuroki

Computer-aided synthesis planning (CASP) algorithms have demonstrated expert-level abilities in planning retrosynthetic routes to molecules of low to moderate complexity. However, current search methods assume the sufficiency of reaching…

Artificial Intelligence · Computer Science 2024-11-04 Kevin Yu , Jihye Roh , Ziang Li , Wenhao Gao , Runzhong Wang , Connor W. Coley

We study off-policy learning (OPL) of contextual bandit policies in large discrete action spaces where existing methods -- most of which rely crucially on reward-regression models or importance-weighted policy gradients -- fail due to…

Machine Learning · Statistics 2024-02-12 Yuta Saito , Jihan Yao , Thorsten Joachims

To address efficiency and design challenges in choice-based matching platforms, we introduce a two-sided assortment optimization framework under general choice preferences. The goal in this problem is to maximize the expected number of…

Optimization and Control · Mathematics 2026-05-08 Omar El Housni , Ulysse Hennebelle , Alfredo Torrico

In Conversational Recommendation Systems (CRS), a user can provide feedback on recommended items at each interaction turn, leading the CRS towards more desirable recommendations. Currently, different types of CRS offer various possibilities…

Information Retrieval · Computer Science 2024-01-12 Maria Vlachou , Craig Macdonald

In fields such as autonomous and safety-critical systems, online optimization plays a crucial role in control and decision-making processes, often requiring the integration of continuous and discrete variables. These tasks are frequently…

Optimization and Control · Mathematics 2025-03-17 Marco Zamponi , Emilio Incerto , Daniele Masti , Mirco Tribastone

Recommender systems generally optimises user engagement, but this approach is dangerous in mental health contexts. When vulnerable users show signs of suicidal ideation, standard algorithms often trap them in echo chambers of harmful…

Information Retrieval · Computer Science 2026-05-26 Alberto Díaz-Álvarez , Raúl Lara-Cabrera , Fernando Ortega-Requena , Víctor Ramos-Osuna

Recommender systems play a key role in shaping modern web ecosystems. These systems alternate between (1) making recommendations (2) collecting user responses to these recommendations, and (3) retraining the recommendation algorithm based…

Information Retrieval · Computer Science 2022-07-18 Karl Krauth , Yixin Wang , Michael I. Jordan

Negative user preference is an important context that is not sufficiently utilized by many existing recommender systems. This context is especially useful in scenarios where the cost of negative items is high for the users. In this work, we…

Information Retrieval · Computer Science 2021-02-19 Bibek Paudel , Sandro Luck , Abraham Bernstein

The voting process is formalized as a multistage voting model with successive alternative elimination. A finite number of agents vote for one of the alternatives each round subject to their preferences. If the number of votes given to the…

Optimization and Control · Mathematics 2018-08-01 Oleg A. Malafeyev , Denis Rylow , Irina Zaitseva , Anna Ermakova , Dmitry Shlaev

A number of attempts have been made to improve accuracy and/or scalability of the PC (Peter and Clark) algorithm, some well known (Buhlmann, et al., 2010; Kalisch and Buhlmann, 2007; 2008; Zhang, 2012, to give some examples). We add here…

Artificial Intelligence · Computer Science 2016-10-06 Joseph Ramsey

The two-time scale nature of SAC, which is an actor-critic algorithm, is characterised by the fact that the critic estimate has not converged for the actor at any given time, but since the critic learns faster than the actor, it ensures…

We study a pessimistic stochastic bilevel program in the context of sequential two-player games, where the leader makes a binary here-and-now decision, and the follower responds a continuous wait-and-see decision after observing the…

Optimization and Control · Mathematics 2022-06-09 Akshit Goyal , Yiling Zhang , Chuan He