中文
相关论文

相关论文: Measuring religious morality using very limited po…

200 篇论文

A goodness-of-fit test for the fitting of a parametric model to data obtained from a detector with finite resolution and limited acceptance is proposed. The parameters of the model are found by minimization of a statistic that is used for…

数据分析、统计与概率 · 物理学 2015-03-17 N. D. Gagunashvili

Much recent work seeks to evaluate values and opinions in large language models (LLMs) using multiple-choice surveys and questionnaires. Most of this work is motivated by concerns around real-world LLM applications. For example,…

In this paper, we demonstrate a data-driven methodology for modelling the local similarity measures of various attributes in a dataset. We analyse the spread in the numerical attributes and estimate their distribution using polynomial…

人工智能 · 计算机科学 2019-05-22 Deepika Verma , Kerstin Bach , Paul Jarle Mork

A sequence of recent papers has considered the role of measurement scales in information retrieval (IR) experimentation, and presented the argument that (only) uniform-step interval scales should be used, and hence that well-known metrics…

信息检索 · 计算机科学 2022-07-08 Alistair Moffat

Psychological assessments commonly rely on rating-scale items, which require respondents to condense complex experiences into predefined categories. Although rich, unstructured text is often captured alongside these scales, it rarely…

计算与语言 · 计算机科学 2026-03-20 Joe Watson , Ivan O'Connor , Chia-Wen Chen , Luning Sun , Fang Luo , David Stillwell

As psychometric surveys are increasingly used to assess the traits of large language models (LLMs), the need for scalable survey item generation suited for LLMs has also grown. A critical challenge here is ensuring the construct validity of…

计算与语言 · 计算机科学 2026-05-26 Sungjib Lim , Woojung Song , Eun-Ju Lee , Yohan Jo

In order to use psychometric instruments to assess a multidimensional construct, we may decompose it in dimensions and, in order to assess each dimension, develop a set of items, so one may assess the construct as a whole, by assessing its…

统计方法学 · 统计学 2018-02-02 Diego Marcondes , Nilton Rogerio Marcondes

The rise of large language models (LLMs) has brought a critical need for high-quality human-labeled data, particularly for processes like human feedback and evaluation. A common practice is to label data via consensus annotation over human…

计算与语言 · 计算机科学 2025-06-23 Manya Wadhwa , Jifan Chen , Junyi Jessy Li , Greg Durrett

Large Language Models (LLMs) are capable of generating opinions and propagating bias unknowingly, originating from unrepresentative and non-diverse data collection. Prior research has analysed these opinions with respect to the West,…

计算机与社会 · 计算机科学 2025-03-11 Hari Shankar , Vedanta S P , Tejas Cavale , Ponnurangam Kumaraguru , Abhijnan Chakraborty

Recently, it was shown that most popular IR measures are not interval-scaled, implying that decades of experimental IR research used potentially improper methods, which may have produced questionable results. However, it was unclear if and…

信息检索 · 计算机科学 2021-01-08 Marco Ferrante , Nicola Ferro , Norbert Fuhr

Previous work has shown that item response theory may be used to rank incorrect response options to multiple-choice items on commonly used assessments. This work has shown that, when the correct response to each item is specified, a nominal…

Opinion dynamics models have been developed to study and predict the evolution of public opinion. Intensive research has been carried out on these models, especially exploring the different rules and topologies, which can be considered two…

物理与社会 · 物理学 2020-10-13 Dino Carpentras , Alejandro Dinkelberg , Michael Quayle

Sensory evaluation is used to assess the consumer acceptance of foods or other consumer products, so as to improve industrial processes and marketing strategies. The procedures currently involved are time-consuming because they require a…

人机交互 · 计算机科学 2019-05-29 M. Guermandi , S. Benatti , D. Brunelli , V. Kartsch , L. Benini

The System Usability Scale (SUS) is a short, survey-based approach used to determine the usability of a system from an end user perspective once a prototype is available for assessment. Individual scores are gathered using a 10-question…

统计方法学 · 统计学 2021-01-26 Nicholas Clark , Matthew Dabkowski , Patrick Driscoll , Dereck Kennedy , Ian Kloo , Heidy Shi

Obtaining quantitative survey responses that are both accurate and informative is crucial to a wide range of fields. Traditional and ubiquitous response formats such as Likert and Visual Analogue Scales require condensation of responses…

人机交互 · 计算机科学 2021-03-11 Zack Ellerby , Christian Wagner , Stephen Broomell

Ensuring the safety of Generative AI requires a nuanced understanding of pluralistic viewpoints. In this paper, we introduce a novel data-driven approach for analyzing ordinal safety ratings in pluralistic settings. Specifically, we address…

Item Response Theory (IRT) models aim to assess latent abilities of $n$ examinees along with latent difficulty characteristics of $m$ test items from categorical data that indicates the quality of their corresponding answers. Classical…

机器学习 · 计算机科学 2024-08-16 Susanne Frick , Amer Krivošija , Alexander Munteanu

Large language models (LLMs) are capable of generating plausible explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can be unfaithful. This,…

计算与语言 · 计算机科学 2025-05-21 Katie Matton , Robert Osazuwa Ness , John Guttag , Emre Kıcıman

This paper analyses changes in public opinion by tracking political discussions in which people voluntarily engage online. Unlike polls or surveys, our approach does not elicit opinions but approximates what the public thinks by analysing…

计算机与社会 · 计算机科学 2010-09-22 Sandra Gonzalez-Bailon , Rafael E. Banchs , Andreas Kaltenbrunner

This paper adapts topic models to the psychometric testing of MOOC students based on their online forum postings. Measurement theory from education and psychology provides statistical models for quantifying a person's attainment of…

机器学习 · 计算机科学 2015-11-26 Jiazhen He , Benjamin I. P. Rubinstein , James Bailey , Rui Zhang , Sandra Milligan , Jeffrey Chan