English
Related papers

Related papers: Evaluating Risks in Weak-to-Strong Alignment: A Bi…

200 papers

A pervasive phenomenon in machine learning applications is distribution shift, where training and deployment conditions for a machine learning model differ. As distribution shift typically results in a degradation in performance, much…

Machine Learning · Statistics 2024-01-23 Philip Amortila , Tongyi Cao , Akshay Krishnamurthy

Interpreting the inference-time behavior of deep neural networks remains a challenging problem. Existing approaches to counterfactual explanation typically ask: What is the closest alternative input that would alter the model's prediction…

Machine Learning · Computer Science 2026-02-12 Brian Hyeongseok Kim , Jacqueline L. Mitchell , Chao Wang

Misclassification detection is an important problem in machine learning, as it allows for the identification of instances where the model's predictions are unreliable. However, conventional uncertainty measures such as Shannon entropy do…

Machine Learning · Statistics 2024-02-09 Eduardo Dadalto , Marco Romanelli , Georg Pichler , Pablo Piantanida

The question-answering (QA) capabilities of foundation models are highly sensitive to prompt variations, rendering their performance susceptible to superficial, non-meaning-altering changes. This vulnerability often stems from the model's…

Machine Learning · Computer Science 2024-06-07 Dyah Adila , Shuai Zhang , Boran Han , Yuyang Wang

In unsupervised ensemble learning, one obtains predictions from multiple sources or classifiers, yet without knowing the reliability and expertise of each source, and with no labeled data to assess it. The task is to combine these possibly…

Machine Learning · Computer Science 2016-02-24 Ariel Jaffe , Ethan Fetaya , Boaz Nadler , Tingting Jiang , Yuval Kluger

In this work, we revisit the weak-to-strong consistency framework, popularized by FixMatch from semi-supervised classification, where the prediction of a weakly perturbed image serves as supervision for its strongly perturbed version.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Lihe Yang , Lei Qi , Litong Feng , Wayne Zhang , Yinghuan Shi

AI for Science (AI4Science) workflows often treat the released dataset as a fixed interface to the underlying system. However, in domains relying on \emph{indirect observation}, the learner observes a derivative representation produced by…

Machine Learning · Computer Science 2026-05-26 Ling Zhan , Xiaoyao Yu , Tao Jia

Deep neural networks are highly susceptible to learning biases in visual data. While various methods have been proposed to mitigate such bias, the majority require explicit knowledge of the biases present in the training data in order to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Rebecca S Stone , Nishant Ravikumar , Andrew J Bulpitt , David C Hogg

The rapid deployment of Large Language Models and AI agents across critical societal and technical domains is hindered by persistent behavioral pathologies including sycophancy, hallucination, and strategic deception that resist mitigation…

Artificial Intelligence · Computer Science 2026-02-23 Xingcheng Xu , Jingjing Qu , Qiaosheng Zhang , Chaochao Lu , Yanqing Yang , Na Zou , Xia Hu

Training data-driven approaches for complex industrial system health monitoring is challenging. When data on faulty conditions are rare or not available, the training has to be performed in a unsupervised manner. In addition, when the…

Machine Learning · Statistics 2021-11-24 Gabriel Michau , Olga Fink

We provide a systematic study of the position-dependent correlation function in weak lensing convergence maps and its relation to the squeezed limit of the three-point correlation function (3PCF) using state-of-the-art numerical…

Cosmology and Nongalactic Astrophysics · Physics 2023-02-14 D. Munshi , G. Jung , T. D. Kitching , J. McEwen , M. Liguori , T. Namikawa , A. Heavens

Reward-model-based fine-tuning is a central paradigm in aligning Large Language Models with human preferences. However, such approaches critically rely on the assumption that proxy reward models accurately reflect intended supervision, a…

Computation and Language · Computer Science 2026-01-21 Zixuan Liu , Siavash H. Khajavi , Guangkai Jiang , Xinru Liu

With the growing accessibility and wide adoption of large language models, concerns about their safety and alignment with human values have become paramount. In this paper, we identify a concerning phenomenon: Reasoning-Induced Misalignment…

Computation and Language · Computer Science 2026-03-11 Hanqi Yan , Hainiu Xu , Siya Qi , Shu Yang , Yulan He

Weakly supervised learning has emerged as a practical alternative to fully supervised learning when complete and accurate labels are costly or infeasible to acquire. However, many existing methods are tailored to specific supervision…

Machine Learning · Computer Science 2025-12-01 Miao Zhang , Junpeng Li , Changchun Hua , Yana Yang

It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenomenon of weak-to-strong generalization exemplifies the advantage of this two-stage procedure:…

Machine Learning · Computer Science 2026-05-26 Diyuan Wu , Lehan Chen , Theodor Misiakiewicz , Marco Mondelli

Student success models might be prone to develop weak spots, i.e., examples hard to accurately classify due to insufficient representation during model creation. This weakness is one of the main factors undermining users' trust, since model…

Machine Learning · Computer Science 2022-12-19 Roberta Galici , Tanja Käser , Gianni Fenu , Mirko Marras

Weak-to-Strong Generalization (W2SG), where a weak model supervises a stronger one, serves as an important analogy for understanding how humans might guide superhuman intelligence in the future. Promising empirical results revealed that a…

Machine Learning · Computer Science 2025-06-19 Yihao Xue , Jiping Li , Baharan Mirzasoleiman

End-to-end models for autonomous driving hold the promise of learning complex behaviors directly from sensor data, but face critical challenges in safety and handling long-tail events. Reinforcement Learning (RL) offers a promising path to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Tianyi Yan , Tao Tang , Xingtai Gui , Yongkang Li , Jiasen Zhesng , Weiyao Huang , Lingdong Kong , Wencheng Han , Xia Zhou , Xueyang Zhang , Yifei Zhan , Kun Zhan , Cheng-zhong Xu , Jianbing Shen

We study reinforcement learning from human feedback under misspecification. Sometimes human feedback is systematically wrong on certain types of inputs, like a broken compass that points the wrong way in specific regions. We prove that when…

Artificial Intelligence · Computer Science 2025-09-16 Madhava Gaikwad

Recommending the best course of action for an individual is a major application of individual-level causal effect estimation. This application is often needed in safety-critical domains such as healthcare, where estimating and communicating…

Machine Learning · Computer Science 2020-10-26 Andrew Jesson , Sören Mindermann , Uri Shalit , Yarin Gal
‹ Prev 1 8 9 10 Next ›