中文
相关论文

相关论文: SimGym: Traffic-Grounded Browser Agents for Offlin…

200 篇论文

Conventional load-testing tools are based on a fifty-year old time-share computer paradigm where a finite number of users submit requests and respond in a synchronized fashion. Conversely, modern web traffic is essentially asynchronous and…

性能 · 计算机科学 2016-09-13 James F. Brady , Neil J. Gunther

Recommender Systems are becoming ubiquitous in many settings and take many forms, from product recommendation in e-commerce stores, to query suggestions in search engines, to friend recommendation in social networks. Current research…

信息检索 · 计算机科学 2018-09-17 David Rohde , Stephen Bonner , Travis Dunlop , Flavian Vasile , Alexandros Karatzoglou

Congestion tollings have been widely developed and adopted as an effective tool to mitigate urban traffic congestion and enhance transportation system sustainability. Nevertheless, these tolling schemes are often tailored on a city-by-city…

应用统计 · 统计学 2024-02-19 Qingnan Liang , Ruili Yao , Ruixuan Zhang , Zhibin Chen , Guoyuan Wu

Alternative recommender systems are critical for ecommerce companies. They guide customers to explore a massive product catalog and assist customers to find the right products among an overwhelming number of options. However, it is a…

信息检索 · 计算机科学 2021-04-16 Mingming Guo , Nian Yan , Xiquan Cui , San He Wu , Unaiza Ahsan , Rebecca West , Khalifeh Al Jadda

Autonomous agents operating in the real world must interact continuously with existing physical and semantic infrastructure, track delayed consequences, and verify outcomes over time. Everyday environments are rich in tangible control…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Jieru Lin , Zhiwei Yu , Börje F. Karlsson

In A/B testing two variants of a piece of software are compared in the field from an end user's point of view, enabling data-driven decision making. While widely used in practice, no comprehensive study has been conducted on the…

软件工程 · 计算机科学 2023-08-10 Federico Quin , Danny Weyns , Matthias Galster , Camila Costa Silva

While rapid advances in large language models (LLMs) are reshaping data-driven intelligent education, accurately simulating students remains an important but challenging bottleneck for scalable educational data collection, evaluation, and…

计算机与社会 · 计算机科学 2025-12-05 Haoxuan Li , Jifan Yu , Xin Cong , Yang Dang , Daniel Zhang-li , Lu Mi , Yisi Zhan , Huiqin Liu , Zhiyuan Liu

Long-horizon planning is widely recognized as a core capability of autonomous LLM-based agents; however, current evaluation frameworks suffer from being largely episodic, domain-specific, or insufficiently grounded in persistent economic…

Online experiments %in which experimental units receive a sequence of treatments over time are frequently employed in many technological companies to evaluate the performance of a newly developed policy, product, or treatment relative to a…

计量经济学 · 经济学 2025-01-14 Ke Sun , Linglong Kong , Hongtu Zhu , Chengchun Shi

Simulating student learning behaviors in open-ended problem-solving environments holds potential for education research, from training adaptive tutoring systems to stress-testing pedagogical interventions. However, collecting authentic data…

人工智能 · 计算机科学 2026-05-07 Hanchen David Wang , Clayton Cohn , Zifan Xu , Siyuan Guo , Gautam Biswas , Meiyi Ma

We study the qualitative and quantitative appearance of stylized facts in several agent-based computational economic market (ABCEM) models. We perform our simulations with the SABCEMM (Simulator for Agent-Based Computational Economic Market…

综合经济学 · 经济学 2019-11-18 Maximilian Beikirch , Simon Cramer , Martin Frank , Philipp Otte , Emma Pabich , Torsten Trimborn

To approach different business objectives, online traffic shaping algorithms aim at improving exposures of a target set of items, such as boosting the growth of new commodities. Generally, these algorithms assume that the utility of each…

机器学习 · 计算机科学 2022-01-03 Chenlin Shen , Guangda Huzhang , Yuhang Zhou , Chen Liang , Qing Da

We present ShoppingComp, a challenging real-world benchmark for comprehensively evaluating LLM-powered shopping agents on three core capabilities: precise product retrieval, expert-level report generation, and safety critical decision…

计算与语言 · 计算机科学 2026-02-10 Huaixiao Tou , Ying Zeng , Yuemeng Li , Cong Ma , Muzhi Li , Minghao Li , Weijie Yuan , He Zhang , Kai Jia

Computer-using agents powered by Vision-Language Models (VLMs) have demonstrated human-like capabilities in operating digital environments like mobile platforms. While these agents hold great promise for advancing digital automation, their…

In this paper, we introduce ECom-Bench, the first benchmark framework for evaluating LLM agent with multimodal capabilities in the e-commerce customer support domain. ECom-Bench features dynamic user simulation based on persona information…

计算与语言 · 计算机科学 2025-11-11 Haoxin Wang , Xianhan Peng , Xucheng Huang , Yizhe Huang , Ming Gong , Chenghan Yang , Yang Liu , Ling Jiang

An increasing number of emerging applications, e.g., internet of things, vehicular communications, augmented reality, and the growing complexity due to the interoperability requirements of these systems, lead to the need to change the tools…

多智能体系统 · 计算机科学 2019-01-16 Merim Dzaferagic , M. Majid Butt , Maria Murphy , Nicholas Kaminski , Nicola Marchetti

We are exploring the enhancement of models of agent behaviour with more "human-like" decision making strategies than are presently available. Our motivation is to developed with a view to as the decision analysis and support for electric…

多智能体系统 · 计算机科学 2009-12-22 Yee Ming Chen , Bo-Yuan Wang , Hung-Ming Shiu

Data science agents promise to accelerate discovery and insight-generation by turning data into executable analyses and findings. Yet existing data science benchmarks fall short due to fragmented evaluation interfaces that make…

人工智能 · 计算机科学 2026-01-26 Fan Nie , Junlin Wang , Harper Hua , Federico Bianchi , Yongchan Kwon , Zhenting Qi , Owen Queen , Shang Zhu , James Zou

The believable simulation of multi-user behavior is crucial for understanding complex social systems. Recently, large language models (LLMs)-based AI agents have made significant progress, enabling them to achieve human-like intelligence…

人工智能 · 计算机科学 2024-12-16 Yijun Liu , Wu Liu , Xiaoyan Gu , Yong Rui , Xiaodong He , Yongdong Zhang

Evaluating multi-turn interactive agents is challenging due to the need for human assessment. Evaluation with simulated users has been introduced as an alternative, however existing approaches typically model generic users and overlook the…

计算与语言 · 计算机科学 2026-03-12 Ryan Shea , Yunan Lu , Liang Qiu , Zhou Yu