中文
相关论文

相关论文: An Online A/B Testing Decision Support System for …

200 篇论文

We describe an investigation of the use of probabilistic models and cost-benefit analyses to guide resource-intensive procedures used by a Web-based question answering system. We first provide an overview of research on question-answering…

信息检索 · 计算机科学 2012-12-12 David Azari , Eric J. Horvitz , Susan Dumais , Eric Brill

Research exploring how to support decision-making has often used machine learning to automate or assist human decisions. We take an alternative approach for improving decision-making, using machine learning to help stakeholders surface ways…

人机交互 · 计算机科学 2023-02-24 Angie Zhang , Olympia Walker , Kaci Nguyen , Jiajun Dai , Anqing Chen , Min Kyung Lee

A/B testing is ubiquitous within the machine learning and data science operations of internet companies. Generically, the idea is to perform a statistical test of the hypothesis that a new feature is better than the existing platform---for…

统计理论 · 数学 2017-10-11 David Goldberg , James E. Johndrow

Decision-making is a fundamental capability in everyday life. Large Language Models (LLMs) provide multifaceted support in enhancing human decision-making processes. However, understanding the influencing factors of LLM-assisted…

人工智能 · 计算机科学 2024-02-28 Eva Eigner , Thorsten Händler

Decision making in crucial applications such as lending, hiring, and college admissions has witnessed increasing use of algorithmic models and techniques as a result of a confluence of factors such as ubiquitous connectivity, ability to…

人工智能 · 计算机科学 2020-09-08 G Roshan Lal , Sahin Cem Geyik , Krishnaram Kenthapadi

As far as Learning Management System is concerned, it offers an integrated platform for educational materials, distribution and management of learning as well as accessibility by a range of users including teachers, learners and content…

人机交互 · 计算机科学 2015-05-12 Selvarajah Thuseethan , Sivapalan Achchuthan , Sinnathamby Kuhanesan

Online experiments (A/B tests) are widely regarded as the gold standard for evaluating recommender system variants and guiding launch decisions. However, a variety of biases can distort the results of the experiment and mislead…

信息检索 · 计算机科学 2025-09-03 Chen Zheng , Zhenyu Zhao

When interpreting A/B tests, we typically focus only on the statistically significant results and take them by face value. This practice, termed post-selection inference in the statistical literature, may negatively affect both point…

应用统计 · 统计学 2021-06-01 Alex Deng , Yicheng Li , Jiannan Lu , Vivek Ramamurthy

Context: User intent modeling is a crucial process in Natural Language Processing that aims to identify the underlying purpose behind a user's request, enabling personalized responses. With a vast array of approaches introduced in the…

We present a methodology to systematically test conversational recommender systems with regards to conversational breakdowns. It involves examining conversations generated between the system and simulated users for a set of pre-defined…

信息检索 · 计算机科学 2024-05-24 Nolwenn Bernard , Krisztian Balog

A/B testing is the foundation of decision-making in online platforms, yet social products often suffer from network interference: user interactions cause treatment effects to spill over into the control group. Such spillovers bias causal…

社会与信息网络 · 计算机科学 2026-02-10 Xu Min , Zhaoxu Yang , Kaixuan Tan , Juan Yan , Xunbin Xiong , Zihao Zhu , Kaiyu Zhu , Fenglin Cui , Yang Yang , Sihua Yang , Jianhui Bu

Despite years of research on testing the usability of mobile applications, our understanding of the issues their users experience still remains fragmented and underexplored. While most earlier studies has provided interesting insights, they…

人机交互 · 计算机科学 2025-12-08 Pawel Weichbroth

As large language models (LLMs) continue to advance, reliable evaluation methods are essential particularly for open-ended, instruction-following tasks. LLM-as-a-Judge enables automatic evaluation using LLMs as evaluators, but its…

计算与语言 · 计算机科学 2025-06-17 Yusuke Yamauchi , Taro Yano , Masafumi Oyamada

Evaluating UX in the context of AI's complexity, unpredictability, and generative nature presents unique challenges. How can we support HCI researchers to create comprehensive UX evaluation plans? In this paper, we introduce EvAlignUX, a…

人机交互 · 计算机科学 2025-07-09 Qingxiao Zheng , Minrui Chen , Pranav Sharma , Yiliu Tang , Mehul Oswal , Yiren Liu , Yun Huang

Machine learning is frequently used in affective computing, but presents challenges due the opacity of state-of-the-art machine learning methods. Because of the impact affective machine learning systems may have on an individual's life, it…

机器学习 · 计算机科学 2025-10-07 David S. Johnson , Olya Hakobyan , Hanna Drimalla

The currently dominating artificial intelligence and machine learning technology, neural networks, builds on inductive statistical learning. Neural networks of today are information processing systems void of understanding and reasoning…

人工智能 · 计算机科学 2022-08-26 Lars Holmberg

Evaluation is crucial in the development process of task-oriented dialogue systems. As an evaluation method, user simulation allows us to tackle issues such as scalability and cost-efficiency, making it a viable choice for large-scale…

信息检索 · 计算机科学 2021-05-11 Weiwei Sun , Shuo Zhang , Krisztian Balog , Zhaochun Ren , Pengjie Ren , Zhumin Chen , Maarten de Rijke

Employees face decisions every day - in the absence of supervision. The outcome of these decisions can be influenced by digital workplace design through the power of persuasive technology. This paper provides a structured literature review…

人机交互 · 计算机科学 2022-08-15 Kilian Wenker

As the importance of comprehensive evaluation in workshop courses increases, there is a growing demand for efficient and fair assessment methods that reduce the workload for faculty members. This paper presents an evaluation conducted with…

计算机与社会 · 计算机科学 2024-05-30 Toru Ishida , Tongxi Liu , Hailong Wang , William K. Cheung

Software qualities such as usability or reliability are among the strongest determinants of mobile app user satisfaction and constitute a significant portion of online user feedback on software products, making it a valuable source of…

软件工程 · 计算机科学 2025-06-16 Eduard C. Groen , Fabiano Dalpiaz , Martijn van Vliet , Boris Winter , Joerg Doerr , Sjaak Brinkkemper