中文
相关论文

相关论文: PredictionMarketBench: A SWE-bench-Style Framework…

200 篇论文

Markets are a promising way to coordinate AI agent activity for similar reasons to those used to justify markets more broadly. In order to effectively participate in markets, agents need to have informative signals of their own ability to…

人工智能 · 计算机科学 2026-04-28 Andrey Fradkin , Rohit Krishnan

Personalized pricing negotiations are a challenging testbed for LLM agents because successful interaction does not guarantee profitable decision making. A seller may produce valid actions and close many deals while still pricing poorly when…

计算机科学与博弈论 · 计算机科学 2026-05-25 Yingjie Lei

This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large Language Model (LLM) agents in the uniquely demanding and fast-paced cryptocurrency domain.…

Large language models (LLMs) demonstrate strong potential as autonomous agents, with promising capabilities in reasoning, tool use, and sequential decision-making. While prior benchmarks have evaluated LLM agents in various domains, the…

机器学习 · 计算机科学 2026-03-03 Yanxu Chen , Zijun Yao , Yantao Liu , Amy Xin , Jin Ye , Jianing Yu , Lei Hou , Juanzi Li

Evaluating AI agents in finance faces two key challenges: static benchmarks require costly expert annotation yet miss the dynamic decision-making central to real-world trading, while LLM-based judges introduce uncontrolled variance on…

人工智能 · 计算机科学 2026-03-03 Xiaochuang Yuan , Hui Xu , Silvia Xu , Cui Zou , Jing Xiong

Large language models (LLMs) achieve strong performance across benchmarks--from knowledge quizzes and math reasoning to web-agent tasks--but these tests occur in static settings, lacking real dynamics and uncertainty. Consequently, they…

交易与市场微观结构 · 定量金融 2025-11-06 Haofei Yu , Fenghai Li , Jiaxuan You

We introduce Prediction Arena, a benchmark for evaluating AI models' predictive accuracy and decision-making by enabling them to trade autonomously on live prediction markets with real capital. Unlike synthetic benchmarks, Prediction Arena…

机器学习 · 计算机科学 2026-04-10 Jaden Zhang , Gardenia Liu , Oliver Johansson , Hileamlak Yitayew , Kamryn Ohly , Grace Li

We introduce MARKET-BENCH, a benchmark that evaluates large language models (LLMs) on introductory quantitative trading tasks by asking them to construct executable backtesters from natural language strategy descriptions and market…

计算与语言 · 计算机科学 2026-01-22 Abhay Srivastava , Sam Jung , Spencer Mateega

Recent advancements have underscored the potential of large language model (LLM)-based agents in financial decision-making. Despite this progress, the field currently encounters two main challenges: (1) the lack of a comprehensive LLM agent…

Predicting real-world events from live market signals demands systems that fuse qualitative news with quantitative order-book dynamics under strict temporal discipline -- a challenge existing benchmarks fail to capture. We present…

计算金融 · 定量金融 2026-04-17 Pu Cheng , Juncheng Liu , Yunshen Long

Prediction markets are powerful mechanisms for information aggregation, but existing designs are optimized for single-event contracts. In practice, traders frequently express beliefs about joint outcomes - through parlays in sports,…

计算工程、金融与科学 · 计算机科学 2026-05-21 Ranvir Rana , Viraj Nadkarni , Niusha Moshrefi , Pramod Viswanath

Prediction markets show considerable promise for developing flexible mechanisms for machine learning. Here, machine learning markets for multivariate systems are defined, and a utility-based framework is established for their analysis. This…

人工智能 · 计算机科学 2015-03-19 Amos Storkey

As Language Models (LMs) increasingly operate as autonomous agents, accurately forecasting their capabilities becomes crucial for societal preparedness. We evaluate six forecasting methods that predict downstream capabilities of LM agents.…

计算与语言 · 计算机科学 2025-03-04 Govind Pimpale , Axel Højmark , Jérémy Scheurer , Marius Hobbhahn

The rapid advancement of Large Language Models (LLMs) has led to a surge of financial benchmarks, evolving from static knowledge evaluation toward interactive trading simulations. However, existing frameworks for evaluating real-time…

交易与市场微观结构 · 定量金融 2026-05-28 Wentao Zhang , Mingxuan Zhao , Jincheng Gao , Jieshun You , Huaiyu Jia , Yilei Zhao , Bo An , Shuo Sun

Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is also a canonical testbed for agentic language models, requiring multi-turn interaction under hidden…

计算机科学与博弈论 · 计算机科学 2026-05-15 Erica Zhang , Fangzhao Zhang , Aneesh Pappu , Batu El , Jose Blanchet , Susan Athey , Jiashuo Liu , James Zou

In this paper, we introduce PredBench, a benchmark tailored for the holistic evaluation of spatio-temporal prediction networks. Despite significant progress in this field, there remains a lack of a standardized framework for a detailed and…

机器学习 · 计算机科学 2024-07-15 ZiDong Wang , Zeyu Lu , Di Huang , Tong He , Xihui Liu , Wanli Ouyang , Lei Bai

Forecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a…

机器学习 · 计算机科学 2025-03-03 Ezra Karger , Houtan Bastani , Chen Yueh-Han , Zachary Jacobs , Danny Halawi , Fred Zhang , Philip E. Tetlock

Prediction markets provide a unique setting where event-level time series are directly tied to natural-language descriptions, yet discovering robust lead-lag relationships remains challenging due to spurious statistical correlations. We…

We introduce TimeSeek, a benchmark for studying how the reliability of agentic LLM forecasters changes over a prediction market's lifecycle. We evaluate 10 frontier models on 150 CFTC-regulated Kalshi binary markets at five temporal…

人工智能 · 计算机科学 2026-04-07 Hamza Mostafa , Om Shastri , Dennis Lee

Large Language Models (LLMs) have achieved impressive results on static code-generation benchmarks, but real-world software development unfolds as a continuous stream of evolving issues, fixes, and feature requests. We introduce…

机器学习 · 计算机科学 2025-07-02 Thomas Joshi , Shayan Chowdhury , Fatih Uysal
‹ 上一页 1 2 3 10 下一页 ›