中文
相关论文

相关论文: Universal Value-Function Uncertainties

200 篇论文

Scaling RL for LLMs is computationally expensive, largely due to multi-sampling for policy optimization and evaluation, making efficient data selection crucial. Inspired by the Zone of Proximal Development (ZPD) theory, we hypothesize LLMs…

机器学习 · 计算机科学 2025-05-20 Yang Zhao , Kai Xiong , Xiao Ding , Li Du , YangouOuyang , Zhouhao Sun , Jiannan Guan , Wenbin Zhang , Bin Liu , Dong Hu , Bing Qin , Ting Liu

Generative Recommendation has emerged as a transformative paradigm, reformulating recommendation as an end-to-end autoregressive sequence generation task. Despite its promise, existing preference optimization methods typically rely on…

信息检索 · 计算机科学 2026-02-13 Chenxiao Fan , Chongming Gao , Yaxin Gong , Haoyan Liu , Fuli Feng , Xiangnan He

Reliable uncertainty quantification in deep neural networks is very crucial in safety-critical applications such as automated driving for trustworthy and informed decision-making. Assessing the quality of uncertainty estimates is…

计算机视觉与模式识别 · 计算机科学 2022-12-12 Neslihan Kose , Ranganath Krishnan , Akash Dhamasia , Omesh Tickoo , Michael Paulitsch

Uncertainty estimation is a key component in any deployed machine learning system. One way to evaluate uncertainty estimation is using "out-of-distribution" (OoD) detection, that is, distinguishing between the training data distribution and…

机器学习 · 计算机科学 2021-12-03 Haiwen Huang , Joost van Amersfoort , Yarin Gal

Uncertainty quantification (UQ) in machine learning is currently drawing increasing research interest, driven by the rapid deployment of deep neural networks across different fields, such as computer vision, natural language processing, and…

机器学习 · 计算机科学 2022-08-26 Zongren Zou , Xuhui Meng , Apostolos F Psaros , George Em Karniadakis

Machine unlearning offers a promising solution to privacy and safety concerns in large language models (LLMs) by selectively removing targeted knowledge while preserving utility. However, current methods are highly sensitive to downstream…

Accurate estimation of intravoxel incoherent motion (IVIM) parameters from diffusion-weighted MRI remains challenging due to the ill-posed nature of the inverse problem and high sensitivity to noise, particularly in the perfusion…

Uncertainty estimation is pivotal in machine learning, especially for classification tasks, as it improves the robustness and reliability of models. We introduce a novel `Epistemic Wrapping' methodology aimed at improving uncertainty…

A set of novel approaches for estimating epistemic uncertainty in deep neural networks with a single forward pass has recently emerged as a valid alternative to Bayesian Neural Networks. On the premise of informative representations, these…

计算机视觉与模式识别 · 计算机科学 2022-07-06 Janis Postels , Mattia Segu , Tao Sun , Luca Sieber , Luc Van Gool , Fisher Yu , Federico Tombari

There have been key advancements to building universal approximators for multi-goal collections of reinforcement learning value functions -- key elements in estimating long-term returns of states in a parameterized manner. We extend this to…

机器学习 · 计算机科学 2024-10-29 Rushiv Arora

Optimal probabilistic approach in reinforcement learning is computationally infeasible. Its simplification consisting in neglecting difference between true environment and its model estimated using limited number of observations causes…

人工智能 · 计算机科学 2013-06-26 Sergey Rodionov , Alexey Potapov , Yurii Vinogradov

Deep hedging trains neural networks to manage derivative risk under market frictions, but produces hedge ratios with no measure of model confidence -- a significant barrier to deployment. We introduce uncertainty quantification to the deep…

计算金融 · 定量金融 2026-03-12 Manan Poddar

Deep neural networks have become the default choice for many of the machine learning tasks such as classification and regression. Dropout, a method commonly used to improve the convergence of deep neural networks, generates an ensemble of…

机器学习 · 统计学 2019-04-11 Tal Kachman , Michal Moshkovitz , Michal Rosen-Zvi

Real-Time Auction (RTA) Interception aims to filter out invalid or irrelevant traffic to enhance the integrity and reliability of downstream data. However, two key challenges remain: (i) the need for accurate estimation of traffic quality…

机器学习 · 计算机科学 2026-05-04 Gaoxiang Zhao , Ruinan Qiu , Pengpeng Zhao , Rongjin Wang , Xiaoting Wang , Zhangang Lin , Xiaoqiang Wang

Model deficiency that results from incomplete training data is a form of structural blindness that leads to costly errors, oftentimes with high confidence. During the training of classification tasks, underrepresented class-conditional…

机器学习 · 计算机科学 2021-02-09 Bruno Abrahao , Zheng Wang , Haider Ahmed , Yuchen Zhu

Automated scientific discovery using large language models relies on identifying genuinely novel solutions. Standard reinforcement learning penalizes high-variance mutations, which leads the policy to prioritize familiar patterns. As a…

Modern vision-based reinforcement learning techniques often use convolutional neural networks (CNN) as universal function approximators to choose which action to take for a given visual input. Until recently, CNNs have been treated like…

机器学习 · 计算机科学 2018-09-28 Jieliang Luo , Sam Green , Peter Feghali , George Legrady , Çetin Kaya Koç

We investigate the fundamental problem of leveraging offline data to accelerate online reinforcement learning - a direction with strong potential but limited theoretical grounding. Our study centers on how to learn and apply value envelopes…

机器学习 · 统计学 2025-10-23 Sebastian Reboul , Hélène Halconruy , Randal Douc

Offline reinforcement learning (RL) provides a promising solution to learning an agent fully relying on a data-driven paradigm. However, constrained by the limited quality of the offline dataset, its performance is often sub-optimal.…

机器学习 · 计算机科学 2026-01-19 Siyuan Guo , Yanchao Sun , Jifeng Hu , Sili Huang , Hechang Chen , Haiyin Piao , Lichao Sun , Yi Chang

Reinforcement Learning (RL) is emerging as tool for tackling complex control and decision-making problems. However, in high-risk environments such as healthcare, manufacturing, automotive or aerospace, it is often challenging to bridge the…

人工智能 · 计算机科学 2022-04-28 Paul Festor , Giulia Luise , Matthieu Komorowski , A. Aldo Faisal
‹ 上一页 1 8 9 10 下一页 ›