中文
相关论文

相关论文: Asymmetric Release Planning-Compromising Satisfact…

200 篇论文

Synchronization among a group of active agents is ubiquitous in nature. Although synchronization based on direct interactions between agents described by the Kuramoto model is well understood, the other general mechanism based on indirect…

统计力学 · 物理学 2025-02-05 Dongliang Zhang , Yuansheng Cao , Qi Ouyang , Yuhai Tu

In this paper, a mathematical negotiation mechanism is designed to minimize the negotiators' costs in a distributed procurement problem at two echelons of an automotive supply chain. The buyer's costs are procurement cost and shortage…

多智能体系统 · 计算机科学 2021-12-21 Zohreh Kaheh , Reza Baradaran Kazemzadeh , Ellips Masehian , Ali Husseinzadeh Kashan

Preference optimization is a critical post-training technique used to align large language models (LLMs) with human preferences, typically by fine-tuning on ranked response pairs. While methods like Direct Preference Optimization (DPO) have…

计算与语言 · 计算机科学 2025-11-12 Rhitabrat Pokharel , Yufei Tao , Ameeta Agrawal

We study sequential multi-issue trading between two greedily rational agents who exchange resources from a finite set of categories. Each agent's utility depends on its allocation, but the offering agent does not know the responding agent's…

多智能体系统 · 计算机科学 2026-05-15 Surya Murthy , Mustafa O. Karabag , Ufuk Topcu

A robust-to-dynamics optimization (RDO) problem is an optimization problem specified by two pieces of input: (i) a mathematical program (an objective function $f:\mathbb{R}^n\rightarrow\mathbb{R}$ and a feasible set…

最优化与控制 · 数学 2023-11-27 Amir Ali Ahmadi , Oktay Gunluk

Several problems in planning and reactive synthesis can be reduced to the analysis of two-player quantitative graph games. {\em Optimization} is one form of analysis. We argue that in many cases it may be better to replace the optimization…

形式语言与自动机理论 · 计算机科学 2021-01-08 Suguman Bansal , Krishnendu Chatterjee , Moshe Y. Vardi

Post-training paradigms for Large Language Models (LLMs), primarily Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), face a fundamental dilemma: SFT provides stability (low variance) but suffers from high fitting bias, while RL…

机器学习 · 计算机科学 2026-04-13 Taojie Zhu , Dongyang Xu , Ding Zou , Sen Zhao , Qiaobo Hao , Zhiguo Yang , Yonghong He

In this book we introduce a new procedure called \alpha-Discounting Method for Multi-Criteria Decision Making (\alpha-D MCDM), which is as an alternative and extension of Saaty Analytical Hierarchy Process (AHP). It works for any number of…

人工智能 · 计算机科学 2015-10-05 Florentin Smarandache

Although Boolean Constraint Technology has made tremendous progress over the last decade, the efficacy of state-of-the-art solvers is known to vary considerably across different types of problem instances and is known to depend strongly on…

人工智能 · 计算机科学 2014-01-07 Holger Hoos , Roland Kaminski , Marius Lindauer , Torsten Schaub

Differential privacy (DP) is the standard for privacy-preserving analysis, and introduces a fundamental trade-off between privacy guarantees and model performance. Selecting the optimal balance is a critical challenge that can be framed as…

机器学习 · 计算机科学 2025-09-05 Yaohong Yang , Aki Rehn , Sammie Katt , Antti Honkela , Samuel Kaski

Reliability-based design optimization (RBDO) is a methodology for designing systems and components under the consideration of probabilistic uncertainty. In practical engineering, the number of input data is often limited, which can damage…

最优化与控制 · 数学 2026-05-27 Takumi Fujiyama , Yoshihiro Kanno

Sequential experiments are often characterized by an exploration-exploitation tradeoff that is captured by the multi-armed bandit (MAB) framework. This framework has been studied and applied, typically when at each time period feedback is…

机器学习 · 计算机科学 2020-12-22 Yonatan Gur , Ahmadreza Momeni

This paper studies distributionally robust regret-optimal (DRRO) control with purified output feedback for linear systems subject to additive disturbances and measurement noise. These uncertainties (including the initial system state) are…

最优化与控制 · 数学 2025-11-21 Shuhao Yan , Carsten W. Scherer

We propose stochastic optimization methodologies for a staffing and capacity planning problem arising from home care practice. Specifically, we consider the perspective of a home care agency that must decide the number of caregivers to hire…

最优化与控制 · 数学 2022-03-29 Ridong Wang , Karmel S. Shehadeh , Xiaolei Xie , Lefei Li

Aligning generative models with human preference via RLHF typically suffers from overoptimization, where an imperfectly learned reward model can misguide the generative model to output undesired responses. We investigate this problem in a…

机器学习 · 计算机科学 2024-12-05 Zhihan Liu , Miao Lu , Shenao Zhang , Boyi Liu , Hongyi Guo , Yingxiang Yang , Jose Blanchet , Zhaoran Wang

We study and extend the semidefinite programming (SDP) hierarchies introduced in [Phys. Rev. Lett. 115, 020501] for the characterization of the statistical correlations arising from finite dimensional quantum systems. First, we introduce…

量子物理 · 物理学 2015-10-28 Miguel Navascues , Adrien Feix , Mateus Araujo , Tamas Vertesi

This paper characterizes the optimal capacity-distortion (C-D) tradeoff in an optical point-to-point system with single-input single-output (SISO) for communication and single-input multiple-output (SIMO) for sensing within an integrated…

信息论 · 计算机科学 2025-04-15 Alireza Ghazavi Khorasgani , Mahtab Mirmohseni , Ahmed Elzanaty

Trustworthy machine learning aims at combating distributional uncertainties in training data distributions compared to population distributions. Typical treatment frameworks include the Bayesian approach, (min-max) distributionally robust…

机器学习 · 计算机科学 2025-05-05 Shixiong Wang , Haowei Wang , Xinke Li , Jean Honorio

We study two-stage distributionally robust optimization (DRO) problems with decision-dependent information discovery (DDID) wherein (a portion of) the uncertain parameters are revealed only if an (often costly) investment is made in the…

最优化与控制 · 数学 2025-10-07 Qing Jin , Angelos Georghiou , Phebe Vayanos , Grani A. Hanasusanto

We study adaptive two-sided assortment optimization for revenue maximization in choice-based matching platforms. The platform has two sides of agents, an initiating side, and a responding side. The decision-maker sequentially selects agents…

计算机科学与博弈论 · 计算机科学 2025-08-13 Mohammadreza Ahmadnejadsaein , Omar El Housni