中文
相关论文

相关论文: Asymmetric Release Planning-Compromising Satisfact…

200 篇论文

Stochastic matching is the stochastic version of the well-known matching problem, which consists in maximizing the rewards of a matching under a set of probability distributions associated with the nodes and edges. In most stochastic…

最优化与控制 · 数学 2024-05-01 Yuya Hikima , Yasunori Akagi , Hideaki Kim

Current alignment pipelines presume a single, universal notion of desirable behavior. However, human preferences often diverge across users, contexts, and cultures. As a result, disagreement collapses into the majority signal and minority…

机器学习 · 计算机科学 2025-06-10 Daniel Halpern , Evi Micha , Ariel D. Procaccia , Itai Shapira

During human motor skill training and physical rehabilitation, there is an inherent trade-off between task difficulty and user performance. Characterizing this trade-off is crucial for evaluating user performance, designing assist-as-needed…

机器人学 · 计算机科学 2026-05-14 Harun Tolasa , Volkan Patoglu

In this work, we study spectrum auction problem where each request from secondary users has spatial, temporal, and spectral features. With the requests of secondary users and the reserve price of the primary user, our goal is to design…

网络与互联网体系结构 · 计算机科学 2013-05-29 Yu-e Sun , He Huang , Xiang-Yang Li , Zhili Chen , Wei Yang , Hongli Xu , Liusheng Huang

Phased releases are a common strategy in the technology industry for gradually releasing new products or updates through a sequence of A/B tests in which the number of treated units gradually grows until full deployment or deprecation.…

机器学习 · 统计学 2023-05-17 Yufan Li , Jialiang Mao , Iavor Bojinov

Stochastic choice-based discrete planning is a broad class of decision-making problems characterized by a sequential decision-making process involving a planner and a group of customers. The firm or planner first decides a subset of options…

最优化与控制 · 数学 2024-09-20 Jiajie Zhang , Yun Hui Lin , Gerardo Berbeglia

Diffusion models have achieved remarkable progress in text-to-image generation, yet aligning them with human preference remains challenging due to the presence of multiple, sometimes conflicting, evaluation metrics (e.g., semantic…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Dipesh Tamboli , Souradip Chakraborty , Aditya Malusare , Biplab Banerjee , Amrit Singh Bedi , Vaneet Aggarwal

Direct alignment methods are increasingly used to align large language models (LLMs) with human preferences. However, many real-world alignment problems involve multiple conflicting objectives, where naive aggregation of preferences can…

计算与语言 · 计算机科学 2026-05-26 Peter Chen , Xiaopeng Li , Xi Chen , Tianyi Lin

We consider an intermediary's problem of dynamically matching demand and supply of heterogeneous types in a periodic-review fashion. More specifically, there are two disjoint sets of demand and supply types, and a reward associated with…

最优化与控制 · 数学 2018-11-20 Ming Hu , Yun Zhou

A matching platform is a system that matches different types of participants, such as companies and job-seekers. In such a platform, merely maximizing the number of matches can result in matches being concentrated on highly popular…

机器学习 · 计算机科学 2026-03-10 Yuki Shibukawa , Koichi Tanaka , Yuta Saito , Shinji Ito

We consider the problem of optimal sparse output feedback controller synthesis for continuous linear time invariant systems when the feedback gain is static and subject to specified structural constraints. Introducing an additional term…

最优化与控制 · 数学 2015-06-23 Reza Arastoo , Nader Motee , Mayuresh V. Kothare

This paper investigates the cooperative planning and control problem for multiple connected autonomous vehicles (CAVs) in different scenarios. In the existing literature, most of the methods suffer from significant problems in computational…

多智能体系统 · 计算机科学 2021-01-05 Xiaoxue Zhang , Zilong Cheng , Jun Ma , Sunan Huang , Frank L. Lewis , Tong Heng Lee

Selecting the appropriate requirements to develop in the next release of an open market software product under evolution, is a compulsory step of each software development project. This selection should be done by maximizing stakeholders'…

软件工程 · 计算机科学 2023-02-07 Jose del Sagrado , Jose Antonio Sierra Ibanez , Isabel M. del Aguila

Achieving optimality and adversarial robustness in deep reinforcement learning has long been regarded as conflicting goals. Nonetheless, recent theoretical insights presented in CAR suggest a potential alignment, raising the important…

机器学习 · 计算机科学 2025-12-02 Haoran Li , Jiayu Lv , Congying Han , Zicheng Zhang , Anqi Li , Yan Liu , Tiande Guo , Nan Jiang

In planning problems, it is often challenging to fully model the desired specifications. In particular, in human-robot interaction, such difficulty may arise due to human's preferences that are either private or complex to model.…

机器人学 · 计算机科学 2021-01-01 Mahsa Ghasemi , Evan Scope Crafts , Bo Zhao , Ufuk Topcu

Direct Preference Optimization (DPO) and its variants have become the de facto standards for aligning large language models (LLMs) with human preferences or specific goals. However, DPO requires high-quality preference data and suffers from…

机器学习 · 计算机科学 2024-11-12 Zhuotong Chen , Fang Liu , Jennifer Zhu , Wanyu Du , Yanjun Qi

Reinforcement learning from human feedback (RLHF) has been extensively employed to align large language models with user intent. However, proximal policy optimization (PPO) based RLHF is occasionally unstable requiring significant…

计算与语言 · 计算机科学 2024-04-02 Saeed Khaki , JinJin Li , Lan Ma , Liu Yang , Prathap Ramachandra

Direct Preference Optimization (DPO) is a simple and efficient framework that has attracted substantial attention. However, it often struggles to meet its primary objectives -- increasing the generation probability of chosen responses while…

人工智能 · 计算机科学 2025-06-17 Jay Hyeon Cho , JunHyeok Oh , Myunsoo Kim , Byung-Jun Lee

Real-world deployments routinely face distribution shifts, group imbalances, and adversarial perturbations, under which the traditional Empirical Risk Minimization (ERM) framework can degrade severely. Distributionally Robust Optimization…

机器学习 · 计算机科学 2026-02-19 Difei Xu , Meng Ding , Zebin Ma , Huanyi Xie , Youming Tao , Aicha Slaitane , Di Wang

Direct Preference Optimization (DPO) has gained significant attention for its simplicity and computational efficiency in aligning large language models (LLMs). Recent advancements have extended DPO to multimodal scenarios, achieving strong…

计算与语言 · 计算机科学 2025-05-27 Yeyuan Wang , Dehong Gao , Rujiao Long , Lei Yi , Linbo Jin , Libin Yang , Xiaoyan Cai