English
Related papers

Related papers: Asymmetric Release Planning-Compromising Satisfact…

200 papers

This study investigates a robust screening problem under distributional ambiguity, where a seller is uncertain about a buyer's true valuation distribution, knowing only that it lies near a reference distribution measured by the Wasserstein…

Optimization and Control · Mathematics 2026-05-19 Shumin Ma , Daniel Zhuoyu Long , Lijian Lu

Preference-based alignment methods (e.g., RLHF, DPO) typically optimize a single scalar objective, implicitly averaging over heterogeneous human preferences. In practice, systematic annotator and user-group disagreement makes mean-reward…

Machine Learning · Computer Science 2026-05-19 Mingxi Zou , Jiaxiang Chen , Junfan Li , Langzhang Liang , Qifan Wang , Xu Yinghui , Zenglin Xu

Improving the alignment of language models with human preferences remains an active research challenge. Previous approaches have primarily utilized Reinforcement Learning from Human Feedback (RLHF) via online RL methods such as Proximal…

Computation and Language · Computer Science 2024-01-25 Tianqi Liu , Yao Zhao , Rishabh Joshi , Misha Khalman , Mohammad Saleh , Peter J. Liu , Jialu Liu

In this paper, we propose a beamforming design that jointly considers two conflicting performance metrics, namely the sum rate and fairness, for a multiple-input single-output non-orthogonal multiple access system. Unlike the conventional…

Signal Processing · Electrical Eng. & Systems 2019-02-18 Haitham Al-Obiedollah , Kanapathippillai Cumanan , Jeyarajan Thiyagalingam , Alister G. Burr , Zhiguo Ding , Octavia A. Dobre

Treatment planning in radiotherapy is inherently a multi-criteria optimization (MCO) problem. Traditionally, the treatment's robustness is not formulated as a part of this decision making problem, but dealt with separately through margins…

Medical Physics · Physics 2026-03-27 Remo Cristoforetti , Philipp Süss , Tobias Becher , Niklas Wahl

The Robust Satisficing (RS) model is an emerging approach to robust optimization, offering streamlined procedures and robust generalization across various applications. However, the statistical theory of RS remains unexplored in the…

Machine Learning · Statistics 2024-06-03 Zhiyi Li , Yunbei Xu , Ruohan Zhan

Multimodal Large Language Models (MLLMs) are powerful at integrating diverse data, but they often struggle with complex reasoning. While Reinforcement learning (RL) can boost reasoning in LLMs, applying it to MLLMs is tricky. Common issues…

Machine Learning · Computer Science 2025-06-30 Minjie Hong , Zirun Guo , Yan Xia , Zehan Wang , Ziang Zhang , Tao Jin , Zhou Zhao

The review-based recommender systems are commonly utilized to measure users preferences towards different items. In this paper, we focus on addressing three main problems existing in the review-based methods. Firstly, these methods suffer…

Information Retrieval · Computer Science 2020-12-14 Yuexin Wu , Tianyu Gao , Sihao Wang , Zhongmin Xiong

Preference handling and optimization are indispensable means for addressing non-trivial applications in Answer Set Programming (ASP). However, their implementation becomes difficult whenever they bring about a significant increase in…

Logic in Computer Science · Computer Science 2011-07-29 Martin Gebser , Roland Kaminski , Torsten Schaub

Decision-makers often encounter uncertainty, and the distribution of uncertain parameters plays a crucial role in making reliable decisions. However, complete information is rarely available. The sample average approximation (SAA) approach…

Optimization and Control · Mathematics 2025-08-27 Ziliang Jin , Jianqiang Cheng , Daniel Zhuoyu Long , Kai Pan

Reviewing the previous work of diversity Rein-forcement Learning,diversity is often obtained via an augmented loss function,which requires a balance between reward and diversity.Generally,diversity optimization algorithms use Multi-armed…

Machine Learning · Computer Science 2024-03-19 Jingcheng Jiang , Haiyin Piao , Yu Fu , Yihang Hao , Chuanlu Jiang , Ziqi Wei , Xin Yang

Aligning large language models (LLMs) with human values is an increasingly critical step in post-training. Direct Preference Optimization (DPO) has emerged as a simple, yet effective alternative to reinforcement learning from human feedback…

Artificial Intelligence · Computer Science 2025-07-29 Yifan Wang , Runjin Chen , Bolian Li , David Cho , Yihe Deng , Ruqi Zhang , Tianlong Chen , Zhangyang Wang , Ananth Grama , Junyuan Hong

Many AI synthesis problems such as planning or scheduling may be modelized as constraint satisfaction problems (CSP). A CSP is typically defined as the problem of finding any consistent labeling for a fixed set of variables satisfying all…

Artificial Intelligence · Computer Science 2013-03-25 Thomas Schiex

Production planning must account for uncertainty in a production system, arising from fluctuating demand forecasts. Therefore, this article focuses on the integration of updated customer demand into the rolling horizon planning cycle. We…

Econometrics · Economics 2024-09-27 Manuel Schlenkrich , Wolfgang Seiringer , Klaus Altendorfer , Sophie N. Parragh

Moment-based distributionally robust optimization (DRO) provides an optimization framework to integrate statistical information with traditional optimization approaches. Under this framework, one assumes that the underlying joint…

Optimization and Control · Mathematics 2023-11-01 Shiyi Jiang , Jianqiang Cheng , Kai Pan , Zuo-Jun Max Shen

Distributionally Robust Optimization (DRO), which aims to find an optimal decision that minimizes the worst case cost over the ambiguity set of probability distribution, has been widely applied in diverse applications, e.g., network…

Machine Learning · Computer Science 2022-12-20 Yang Jiao , Kai Yang , Dongjin Song

Remanufacturing is pivotal in transitioning to more sustainable economies. While industry evidence highlights its vast market potential and economic and environmental benefits, remanufacturing remains underexplored in theoretical research.…

Optimization and Control · Mathematics 2025-01-07 Amirreza Pashapour , Fatemeh Zare Bidaki

We consider a discrete-time bipartite matching model with random arrivals of units of supply and demand that can wait in queues located at the nodes in the network. A control policy determines which are matched at each time. The focus is on…

Discrete Mathematics · Computer Science 2016-06-28 Ana Bušić , Sean Meyn

The field of preference optimization has made outstanding contributions to the alignment of language models with human preferences. Despite these advancements, recent methods still rely heavily on substantial paired (labeled) feedback data,…

Machine Learning · Computer Science 2026-02-20 Seonggyun Lee , Sungjun Lim , Seojin Park , Soeun Cheon , Kyungwoo Song

This study introduces adaptive robust optimization (ARO) and adaptive robust stochastic optimization (ARSO) approaches to address long- and short-term uncertainties in the optimal sizing and placement of distributed energy resources in…

Optimization and Control · Mathematics 2025-03-25 Fernando García-Muñoz , Cristian Duran-Mateluna