中文
相关论文

相关论文: IRIS: Interpretable Rubric-Informed Segmentation f…

200 篇论文

Understanding an agent's goals from its behavior is fundamental to aligning AI systems with human intentions. Existing goal recognition methods typically rely on an optimal goal-oriented policy representation, which may differ from the…

人工智能 · 计算机科学 2026-02-17 Osher Elhadad , Felipe Meneguzzi , Reuth Mirsky

With the recent advances in A.I. methodologies and their application to medical imaging, there has been an explosion of related research programs utilizing these techniques to produce state-of-the-art classification performance. Ultimately,…

Assessing action quality is both imperative and challenging due to its significant impact on the quality of AI-generated videos, further complicated by the inherently ambiguous nature of actions within AI-generated video (AIGV). Current…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Zijian Chen , Wei Sun , Yuan Tian , Jun Jia , Zicheng Zhang , Jiarui Wang , Ru Huang , Xiongkuo Min , Guangtao Zhai , Wenjun Zhang

Assessing action quality is challenging due to the subtle differences between videos and large variations in scores. Most existing approaches tackle this problem by regressing a quality score from a single video, suffering a lot from the…

计算机视觉与模式识别 · 计算机科学 2021-08-18 Xumin Yu , Yongming Rao , Wenliang Zhao , Jiwen Lu , Jie Zhou

Evaluating the abilities of learners is a fundamental objective in the field of education. In particular, there is an increasing need to assess higher-order abilities such as expressive skills and logical thinking. Constructed-response…

计算与语言 · 计算机科学 2025-06-26 Masaki Uto , Yuma Ito

Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, handwritten single-variable calculus work from UC Irvine. Using…

机器学习 · 计算机科学 2026-03-03 Zhiqi Yu , Xingping Liu , Haobin Mao , Mingshuo Liu , Long Chen , Jack Xin , Yifeng Yu

We propose Partially Interpretable Estimators (PIE) which attribute a prediction to individual features via an interpretable model, while a (possibly) small part of the PIE prediction is attributed to the interaction of features via a…

机器学习 · 计算机科学 2021-05-07 Tong Wang , Jingyi Yang , Yunyi Li , Boxiang Wang

Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with human scoring standards remains challenging. We formulate this challenge as a criteria-transfer…

计算与语言 · 计算机科学 2026-05-29 Yihan Hong , Huaiyuan Yao , Bolin Shen , Wanpeng Xu , Hua Wei , Yushun Dong

LLM-based automated scoring approaches near-human performance, but scaling to new tasks remains bottlenecked by the per-item human configuration of upstream stages such as rubric construction. Human experts bypass this bottleneck through…

计算与语言 · 计算机科学 2026-05-29 Yun Wang , Xin Xia , Xuansheng Wu , Xiaoming Zhai , Ninghao Liu

Visual Quality Assessment (QA) seeks to predict human perceptual judgments of visual fidelity. While recent multimodal large language models (MLLMs) show promise in reasoning about image and video quality, existing approaches mainly rely on…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Zehui Feng , Tian Qiu , Tong Wu , Junxuan Li , Huayuan Xu , Ting Han

Video Question Answering (VideoQA) aims to answer natural language questions based on the given video, with prior work primarily focusing on identifying the duration of relevant segments, referred to as explicit visual evidence. However,…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Tieyuan Chen , Huabin Liu , Yi Wang , Chaofan Gan , Mingxi Lyu , Ziran Qin , Shijie Li , Liquan Shen , Junhui Hou , Zheng Wang , Weiyao Lin

End-of-course assessments play important roles in the ongoing attempt to improve instruction in physics courses. Comparison of students' performance on assessments before and after instruction gives a measure of student learning. In…

物理教育 · 物理学 2014-07-15 Leanne Doughty , Marcos D. Caballero

Sentiment Analysis Systems (SASs) are data-driven Artificial Intelligence (AI) systems that, given a piece of text, assign one or more numbers conveying the polarity and emotional intensity expressed in the input. Like other automatic…

人工智能 · 计算机科学 2023-02-07 Kausik Lakkaraju , Biplav Srivastava , Marco Valtorta

Risk scoring systems have been widely deployed in many applications, which assign risk scores to users according to their behavior sequences. Though many deep learning methods with sophisticated designs have achieved promising results, the…

机器学习 · 计算机科学 2022-08-17 Yao Zhang , Yun Xiong , Yiheng Sun , Caihua Shan , Tian Lu , Hui Song , Yangyong Zhu

We introduce SIRUS (Stable and Interpretable RUle Set) for regression, a stable rule learning algorithm which takes the form of a short and simple list of rules. State-of-the-art learning algorithms are often referred to as "black boxes"…

机器学习 · 统计学 2021-02-11 Clément Bénard , Gérard Biau , Sébastien da Veiga , Erwan Scornet

Action Quality Assessment (AQA) predicts fine-grained execution scores from action videos and is widely applied in sports, rehabilitation, and skill evaluation. Long-term AQA, as in figure skating or rhythmic gymnastics, is especially…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Ruisheng Han , Kanglei Zhou , Shuang Chen , Amir Atapour-Abarghouei , Hubert P. H. Shum

Assessing fairness in artificial intelligence (AI) typically involves AI experts who select protected features, fairness metrics, and set fairness thresholds to assess outcome fairness. However, little is known about how stakeholders,…

人工智能 · 计算机科学 2026-02-27 Lin Luo , Yuri Nakao , Mathieu Chollet , Hiroya Inakoshi , Simone Stumpf

Video quality assessment (VQA) is a fundamental computer vision task that aims to predict the perceptual quality of a given video in alignment with human judgments. Existing performant VQA models trained with direct score supervision suffer…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Shuo Xing , Soumik Dey , Mingyang Wu , Ashirbad Mishra , Naveen Ravipati , Binbin Li , Hansi Wu , Zhengzhong Tu

In many real-world continuous action domains, human agents must decide which actions to attempt and then execute those actions to the best of their ability. However, humans cannot execute actions without error. Human performance in these…

人工智能 · 计算机科学 2024-08-21 Delma Nieves-Rivera , Christopher Archibald

Evaluating fairness under domain shift is challenging because scalar metrics often obscure exactly where and how disparities arise. We introduce \textit{RISE} (Residual Inspection through Sorted Evaluation), an interactive visualization…

机器学习 · 计算机科学 2026-02-05 Ray Chen , Christan Grant