English
Related papers

Related papers: Human-Centric Evaluation for Foundation Models

200 papers

Recent advancements in foundation models have improved autonomous tool usage and reasoning, but their capabilities in map-based reasoning remain underexplored. To address this, we introduce MapEval, a benchmark designed to assess foundation…

DeepSeek-V3 and DeepSeek-R1 are leading open-source Large Language Models (LLMs) for general-purpose tasks and reasoning, achieving performance comparable to state-of-the-art closed-source models from companies like OpenAI and Anthropic --…

Machine Learning · Computer Science 2025-03-17 Chengen Wang , Murat Kantarcioglu

Supportive conversation depends on skills that go beyond language fluency, including reading emotions, adjusting tone, and navigating moments of resistance, frustration, or distress. Despite rapid progress in language models, we still lack…

Computation and Language · Computer Science 2026-02-26 Laya Iyer , Kriti Aggarwal , Sanmi Koyejo , Gail Heyman , Desmond C. Ong , Subhabrata Mukherjee

Person re-identification (Person ReID) is a challenging task due to the large variations in camera viewpoint, lighting, resolution, and human pose. Recently, with the advancement of deep learning technologies, the performance of Person ReID…

Computer Vision and Pattern Recognition · Computer Science 2018-05-02 Wangmeng Xiang , Jianqiang Huang , Xianbiao Qi , Xiansheng Hua , Lei Zhang

Generative Search Engines (GSEs) synthesize conversational answers from multiple sources, weakening the long-standing link between search ranking and digital visibility. This shift raises a central question for content creators: How can we…

Computation and Language · Computer Science 2025-12-29 Qiyuan Chen , Jiahe Chen , Hongsen Huang , Qian Shao , Jintai Chen , Renjie Hua , Hongxia Xu , Ruijia Wu , Ren Chuan , Jian Wu

Humans develop perception through a bottom-up hierarchy: from basic primitives and Gestalt principles to high-level semantics. In contrast, current Multimodal Large Language Models (MLLMs) are trained directly on complex downstream tasks,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Jen-Tse Huang , Dasen Dai , Jen-Yuan Huang , Youliang Yuan , Xiaoyuan Liu , Wenxuan Wang , Wenxiang Jiao , Pinjia He , Zhaopeng Tu , Haodong Duan

Most frontier AI developers publicly document their safety evaluations of new AI models in model reports, including testing for chemical and biological (ChemBio) misuse risks. This practice provides a window into the methodology of these…

Computers and Society · Computer Science 2025-10-29 Tom Reed , Tegan McCaslin , Luca Righetti

This entry provides an overview of Human-centered Geospatial Data Science, highlighting the gaps it aims to bridge, its significance, and its key topics and research. Geospatial Data Science, which derives geographic knowledge and insights…

Computers and Society · Computer Science 2025-01-13 Yuhao Kang

Human evaluation of machine translation normally uses sentence-level measures such as relative ranking or adequacy scales. However, these provide no insight into possible errors, and do not scale well with sentence length. We argue for a…

Computation and Language · Computer Science 2016-09-28 Alexandra Birch , Omri Abend , Ondrej Bojar , Barry Haddow

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Yi Yuan , Jingdong Chen , Le Wang

Multimodal models have achieved remarkable progress in recent years. Nevertheless, they continue to exhibit notable limitations in spatial understanding and reasoning, the very capability that anchors artificial general intelligence in the…

In the past decade, deep learning (DL) models have gained prominence for their exceptional accuracy on benchmark datasets in recommender systems (RecSys). However, their evaluation has primarily relied on offline metrics, overlooking direct…

Information Retrieval · Computer Science 2024-05-03 Ruixuan Sun , Xinyi Wu , Avinash Akella , Ruoyan Kong , Bart Knijnenburg , Joseph A. Konstan

A personalized LLM should remember user facts, apply them correctly, and adapt over time to provide responses that the user prefers. Existing LLM personalization benchmarks are largely centered on two axes: accurately recalling user…

Machine Learning · Computer Science 2025-12-16 Md Awsafur Rahman , Adam Gabrys , Doug Kang , Jingjing Sun , Tian Tan , Ashwin Chandramouli

Egocentric human-object interaction (Ego-HOI) detection is crucial for intelligent agents to understand and assist human activities from a first-person perspective. However, progress has been hindered by the lack of benchmarks and methods…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Kunyuan Deng , Yi Wang , Lap-Pui Chau

We introduce HonestCyberEval, a new benchmark for assessing AI models' capabilities and risks in automated software exploitation, focusing on their ability to detect and exploit vulnerabilities in real-world software systems. Our evaluation…

Cryptography and Security · Computer Science 2025-08-27 Dan Ristea , Vasilios Mavroudis

The human capital invested into software development plays a vital role in the success of any software project. By human capital, we do not mean the individuals themselves, but involves the range of knowledge and skills (i.e., human…

Software Engineering · Computer Science 2018-05-11 Saya Onoue , Hideaki Hata , Raula Gaikovina Kula , Kenichi Matsumoto

AI measurement science has a wide variety of methodologies and measurements for comparing AI systems, resulting in what often appear to be "apples-to-oranges" comparisons across AI evaluations. To move toward "apples-to-apples" comparisons…

Human-Computer Interaction · Computer Science 2026-05-11 Yee-Yin Choong , Kristen Greene , Alice Qian , Meryem Marasli , Ziqi Yang , Sophia Chen , Laura Dabbish , Anand Rao , Hong Shen

Our work focuses on the social reasoning capabilities of foundation models for real-world human-robot interactions. We introduce the Social Human Robot Embodied Conversation (SHREC) Dataset, a benchmark of $\sim$400 real-world human-robot…

Human-Computer Interaction · Computer Science 2026-05-13 Dong Won Lee , Yubin Kim , Denison Guvenoz , Sooyeon Jeong , Parker Malachowsky , Louis-Philippe Morency , Cynthia Breazeal , Hae Won Park

As dialogue systems and chatbots increasingly integrate into everyday interactions, the need for efficient and accurate evaluation methods becomes paramount. This study explores the comparative performance of human and AI assessments across…

Computation and Language · Computer Science 2024-09-11 Ike Ebubechukwu , Johane Takeuchi , Antonello Ceravola , Frank Joublin

Geometric problem solving constitutes a critical branch of mathematical reasoning, requiring precise analysis of shapes and spatial relationships. Current evaluations of geometric reasoning in vision-language models (VLMs) face limitations,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Yuan Feng , Yue Yang , Xiaohan He , Jiatong Zhao , Jianlong Chen , Zijun Chen , Daocheng Fu , Qi Liu , Renqiu Xia , Bo Zhang , Junchi Yan