中文
相关论文

相关论文: What Do We Actually Learn from Evaluations in the …

200 篇论文

Evaluation is the foundation of empirical science, yet the evaluation of evaluation itself -- so-called meta-evaluation -- remains strikingly underdeveloped. While methods such as observational studies, design of experiments (DoE), and…

统计方法学 · 统计学 2026-01-22 Hongxiao Li , Chenxi Wang , Fanda Fan , Zihan Wang , Wanling Gao , Lei Wang , Jianfeng Zhan

In contrast to objectively measurable aspects (such as accuracy, reading speed, or memorability), the subjective experience of visualizations has only recently gained importance, and we have less experience how to measure it. We explore how…

人机交互 · 计算机科学 2023-10-24 Laura Koesten , Drew Dimmery , Michael Gleicher , Torsten Möller

Visualization design influences how people perceive data patterns, yet most research focuses on low-level analytic tasks, such as finding correlations. The extent to which these perceptual affordances translate to high-level decision-making…

人机交互 · 计算机科学 2025-05-09 Yixuan Li , Emery D. Berger , Minsuk Kahng , Cindy Xiong Bearfield

There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluation ecosystem is insufficient: Commonly used static benchmarks…

More visualization systems are simplifying the data analysis process by automatically suggesting relevant visualizations. However, little work has been done to understand if users trust these automated recommendations. In this paper, we…

人机交互 · 计算机科学 2021-04-07 Rachael Zehrung , Astha Singhal , Michael Correll , Leilani Battle

Automated approaches to answer patient-posed health questions are rising, but selecting among systems requires reliable evaluation. The current gold standard for evaluating the free-text artificial intelligence (AI) responses--human expert…

人工智能 · 计算机科学 2026-05-11 Sarvesh Soni , Dina Demner-Fushman

In this paper we present an overview of several visualization techniques to support the search process in Digital Libraries (DLs). The search process typically can be separated into three major phases: query formulation and refinement,…

数字图书馆 · 计算机科学 2013-04-16 Wilko van Hoek , Philipp Mayr

Various standardized tests exist that assess individuals' visualization literacy. Their use can help to draw conclusions from studies. However, it is not taken into account that the test itself can create a pressure situation where…

人机交互 · 计算机科学 2024-09-13 Seyda Öney , Moataz Abdelaal , Kuno Kurzhals , Paul Betz , Cordula Kropp , Daniel Weiskopf

How well can large language models (LLMs) generate summaries? We develop new datasets and conduct human evaluation experiments to evaluate the zero-shot generation capability of LLMs across five distinct summarization tasks. Our findings…

计算与语言 · 计算机科学 2023-09-19 Xiao Pu , Mingqi Gao , Xiaojun Wan

As AI systems advance beyond human capabilities, scalable oversight becomes critical: how can we supervise AI that exceeds our abilities? A key challenge is that human evaluators may form incorrect beliefs about AI behavior in complex…

人工智能 · 计算机科学 2025-10-22 Leon Lang , Patrick Forré

It is widely believed that theory is useful in physics because it describes simple systems and that strictly empirical phenomenological approaches are necessary for complex biological and social systems. Here we prove based upon an analysis…

物理与社会 · 物理学 2013-08-15 Yaneer Bar-Yam

Human evaluation is critical for validating the performance of text-to-image generative models, as this highly cognitive process requires deep comprehension of text and images. However, our survey of 37 recent papers reveals that many works…

计算机视觉与模式识别 · 计算机科学 2023-04-05 Mayu Otani , Riku Togashi , Yu Sawai , Ryosuke Ishigami , Yuta Nakashima , Esa Rahtu , Janne Heikkilä , Shin'ichi Satoh

Time series forecasting is essential for agents to make decisions. Traditional approaches rely on statistical methods to forecast given past numeric values. In practice, end-users often rely on visualizations such as charts and plots to…

计算机视觉与模式识别 · 计算机科学 2021-11-23 Srijan Sood , Zhen Zeng , Naftali Cohen , Tucker Balch , Manuela Veloso

Driven by vast and diverse textual data, large language models (LLMs) have demonstrated impressive performance across numerous natural language processing (NLP) tasks. Yet, a critical question persists: does their generalization arise from…

计算与语言 · 计算机科学 2025-09-08 Boxiang Ma , Ru Li , Yuanlong Wang , Hongye Tan , Xiaoli Li

As visualization researchers evaluate the impact of visualization design on decision-making, they often hold a one-dimensional perspective on the cognitive processes behind making a decision. Several psychological and economical researchers…

人机交互 · 计算机科学 2020-10-09 Melanie Bancilhon , Alvitta Ottley

Neural language models (LMs) are arguably less data-efficient than humans from a language acquisition perspective. One fundamental question is why this human-LM gap arises. This study explores the advantage of grounded language acquisition,…

计算与语言 · 计算机科学 2024-12-18 Tatsuki Kuribayashi , Timothy Baldwin

Pre-trained language models are still far from human performance in tasks that need understanding of properties (e.g. appearance, measurable quantity) and affordances of everyday objects in the real world since the text lacks such…

计算与语言 · 计算机科学 2022-03-18 Woojeong Jin , Dong-Ho Lee , Chenguang Zhu , Jay Pujara , Xiang Ren

With the rise of machines to human-level performance in complex recognition tasks, a growing amount of work is directed towards comparing information processing in humans and machines. These studies are an exciting chance to learn about one…

计算机视觉与模式识别 · 计算机科学 2021-04-14 Christina M. Funke , Judy Borowski , Karolina Stosio , Wieland Brendel , Thomas S. A. Wallis , Matthias Bethge

Clinician-facing predictive models are increasingly present in the healthcare setting. Regardless of their success with respect to performance metrics, all models have uncertainty. We investigate how to visually communicate uncertainty in…

人机交互 · 计算机科学 2022-10-25 Caitlin F. Harrigan , Gabriela Morgenshtern , Anna Goldenberg , Fanny Chevalier

The ability of to explain neural network decisions goes hand in hand with their safe deployment. Several methods have been proposed to highlight features important for a given network decision. However, there is no consensus on how to…

计算机视觉与模式识别 · 计算机科学 2020-03-23 Agnieszka Grabska-Barwińska
‹ 上一页 1 8 9 10 下一页 ›