English
Related papers

Related papers: Application of accelerated life testing in human r…

200 papers

In order to interpret and explain the physiological signal behaviors, it can be interesting to find some constants among the fluctuations of these data during all the effort or during different stages of the race (which can be detected…

Applications · Statistics 2011-12-06 Imen Kammoun , Véronique Billat , Jean-Marc Bardet

While Reinforcement Learning from Human Feedback (RLHF) is widely used to align Large Language Models (LLMs) with human preferences, it typically assumes homogeneous preferences across users, overlooking diverse human values and minority…

Computation and Language · Computer Science 2025-10-28 Yijiang River Dong , Tiancheng Hu , Yinhong Liu , Ahmet Üstün , Nigel Collier

Harmful fine-tuning (HFT), performed directly on open-source LLMs or through Fine-tuning-as-a-Service, breaks safety alignment and poses significant threats. Existing methods aim to mitigate HFT risks by learning robust representation on…

Machine Learning · Computer Science 2025-08-13 Liang Chen , Xueting Han , Li Shen , Jing Bai , Kam-Fai Wong

Fatigue data arise in many research and applied areas and there have been statistical methods developed to model and analyze such data. The distributions of fatigue life and fatigue strength are often of interest to engineers designing…

Methodology · Statistics 2024-03-20 Peng Liu , Yili Hong , Luis A. Escobar , William Q. Meeker

Large vision-language models (VLMs) achieve strong performance in Visual Question Answering but still rely heavily on supervised fine-tuning (SFT) with massive labeled datasets, which is costly due to human annotations. Crucially,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Jian Lan , Zhicheng Liu , Udo Schlegel , Raoyuan Zhao , Yihong Liu , Hinrich Schütze , Michael A. Hedderich , Thomas Seidl

We introduce adaptive learn-then-test (aLTT), an efficient hyperparameter selection procedure that provides finite-sample statistical guarantees on the population risk of AI models. Unlike the existing learn-then-test (LTT) technique, which…

Machine Learning · Statistics 2025-02-03 Matteo Zecchin , Sangwoo Park , Osvaldo Simeone

Supervisory-based human-robot teams are deployed in various dynamic and extreme environments (e.g., space exploration). Achieving high task performance in such environments is critical, as a mistake may lead to significant monetary loss or…

Robotics · Computer Science 2020-03-13 Jamison Heard , Julian Fortune , Julie A. Adams

Measuring user satisfaction level is a challenging task, and a critical component in developing large-scale conversational agent systems serving the needs of real users. An widely used approach to tackle this is to collect human annotation…

We study \emph{Human Projection} (HP): people's tendency to evaluate AI using the same frameworks they use for humans -- treating features such as task difficulty and the reasonableness of mistakes as diagnostic of overall ability. We…

General Economics · Economics 2026-05-12 Bnaya Dreyfuss , Raphaël Raux

The effectiveness of automatic evaluation of generative models is typically measured by comparing the labels generated via automation with labels by humans using correlation metrics. However, metrics like Krippendorff's $\alpha$ and…

Human-Computer Interaction · Computer Science 2025-01-28 Aparna Elangovan , Lei Xu , Jongwoo Ko , Mahsa Elyasi , Ling Liu , Sravan Bodapati , Dan Roth

The paper discusses the challenge of evaluating the prognosis quality of machine health index (HI) data. Many existing solutions in machine health forecasting involve visually assessing the quality of predictions to roughly gauge the…

Signal Processing · Electrical Eng. & Systems 2025-02-14 Daniel Kuzio , Radosław Zimroz , Agnieszka Wyłomańska

The work deals with the fatigue lifetime estimation of Short Fiber Reinforced Thermoplastics (SFRP), with a focus on conjugated effects of thermal aging. Two materials containing 35% (V35) and 50% (V50) weight ratio of short glass fibers…

Materials Science · Physics 2022-11-18 Florent Alexis , Sylvie Castagnet , Carole Nadot-Martin , Gilles Robert , Peggy Havet

Validation of autonomous driving systems requires a trade-off between test fidelity, cost, and scalability. While miniaturized hardware-in-the-loop (HIL) platforms have emerged as a promising solution, a systematic framework supporting…

Robotics · Computer Science 2025-10-22 Mingxin Li , Haibo Hu , Jinghuai Deng , Yuchen Xi , Xinhong Chen , Jianping Wang

Reinforcement learning from human feedback (RLHF) is fundamentally limited by the capacity of humans to correctly evaluate model output. To improve human evaluation ability and overcome that limitation this work trains "critic" models that…

Software Engineering · Computer Science 2024-07-02 Nat McAleese , Rai Michael Pokorny , Juan Felipe Ceron Uribe , Evgenia Nitishinskaya , Maja Trebacz , Jan Leike

Existing test-time adaptation (TTA) approaches often adapt models with the unlabeled testing data stream. A recent attempt relaxed the assumption by introducing limited human annotation, referred to as Human-In-the-Loop Test-Time Adaptation…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Yushu Li , Yongyi Su , Xulei Yang , Kui Jia , Xun Xu

Aligning human preference and value is an important requirement for building contemporary foundation models and embodied AI. However, popular approaches such as reinforcement learning with human feedback (RLHF) break down the task into…

Artificial Intelligence · Computer Science 2024-12-03 Chenliang Li , Siliang Zeng , Zeyi Liao , Jiaxiang Li , Dongyeop Kang , Alfredo Garcia , Mingyi Hong

Reinforcement Learning (RL) can be extremely effective in solving complex, real-world problems. However, injecting human knowledge into an RL agent may require extensive effort and expertise on the human designer's part. To date, human…

Artificial Intelligence · Computer Science 2018-05-16 Ariel Rosenfeld , Moshe Cohen , Matthew E. Taylor , Sarit Kraus

Health Indicators (HIs) are essential for predicting system failures in predictive maintenance. While methods like RaPP (Reconstruction along Projected Pathways) improve traditional HI approaches by leveraging autoencoder latent spaces,…

Performance · Computer Science 2025-07-10 Lucas Thil , Jesse Read , Rim Kaddah , Guillaume Florent Doquet

Re-inforcement learning from human feedback (RLHF) has been effective in the task of AI alignment. However, one of the key assumptions of RLHF is that the annotators (referred to as workers from here on out) have a homogeneous response…

Human-Computer Interaction · Computer Science 2026-01-29 Sarvesh Shashidhar , Abhishek Mishra , Madhav Kotecha

Research on Automatic Story Generation (ASG) relies heavily on human and automatic evaluation. However, there is no consensus on which human evaluation criteria to use, and no analysis of how well automatic criteria correlate with them. In…

Computation and Language · Computer Science 2022-09-16 Cyril Chhun , Pierre Colombo , Chloé Clavel , Fabian M. Suchanek