中文
相关论文

相关论文: PredictionMarketBench: A SWE-bench-Style Framework…

200 篇论文

Predictive Process Analytics is becoming an essential aid for organizations, providing online operational support of their processes. However, process stakeholders need to be provided with an explanation of the reasons why a given process…

Current benchmarks for evaluating software engineering agents, such as SWE-Bench Verified, are predominantly derived from GitHub issues and fail to accurately reflect how developers interact with chat-based coding assistants in integrated…

软件工程 · 计算机科学 2026-01-27 Spandan Garg , Benjamin Steenhoek , Yufan Huang

Real-world financial decision-making is a challenging problem that requires reasoning over heterogeneous signals, including company fundamentals derived from regulatory filings and trading signals computed from price dynamics. Recently,…

计算工程、金融与科学 · 计算机科学 2026-03-24 Yogesh Agrawal , Aniruddha Dutta , Md Mahadi Hasan , Santu Karmaker , Aritra Dutta

We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering. To this end, we curate 75 ML engineering-related competitions from Kaggle, creating a diverse set of challenging tasks that test…

Predictive models that are developed in a regulated industry or a regulated application, like determination of credit worthiness, must be interpretable and rational (e.g., meaningful improvements in basic credit behavior must result in…

机器学习 · 统计学 2018-06-13 Bob Vanderheyden , Jennifer Priestley

Financial risk prediction plays a crucial role in the financial sector. Machine learning methods have been widely applied for automatically detecting potential risks and thus saving the cost of labor. However, the development in this field…

风险管理 · 定量金融 2023-08-02 Yuwei Yin , Yazheng Yang , Jian Yang , Qi Liu

We propose a dynamic model of a prediction market in which agents predict the values of a sequence of random vectors. The main result shows that if there are agents who make correct (or asymptotically correct) next-period forecasts, then…

概率论 · 数学 2024-02-27 Nina Badulina , Dmitry Shatilovich , Mikhail Zhitlukhin

In the realm of cryptocurrency, the prediction of Bitcoin prices has garnered substantial attention due to its potential impact on financial markets and investment strategies. This paper propose a comparative study on hybrid machine…

机器学习 · 计算机科学 2024-01-02 Shun Liu , Kexin Wu , Chufeng Jiang , Bin Huang , Danqing Ma

Prediction markets are used in real life to predict outcomes of interest such as presidential elections. This paper presents a mathematical theory of artificial prediction markets for supervised learning of conditional probability…

机器学习 · 统计学 2015-03-18 Adrian Barbu , Nathan Lay

Daily probability changes in Kalshi macro prediction markets forecast cryptocurrency realized volatility through two distinct channels. The monetary policy channel, measured by Fed rate repricing on KXFED contracts, predicts Bitcoin…

统计金融 · 定量金融 2026-04-03 Hardhik Mohanty , Bhaskar Krishnamachari

Reinforcement learning agents for portfolio management are typically trained and deployed as static policies, with no mechanism for using price forecasts at inference time. We propose $\text{FPILOT}$ (**Fin**ancial **P**lugin…

机器学习 · 计算机科学 2026-05-14 Eun Go , Rohan Deb , Arindam Banerjee

As agentic network management gains popularity, there is a critical need for evaluation frameworks that transcend static, one-shot testing. To address this, we introduce NetAgentBench, a dynamic benchmark that evaluates agent interactions…

网络与互联网体系结构 · 计算机科学 2026-04-14 Ahmed Twabi , Yepeng Ding , Tohru Kondo

While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reasoning where errors are often rectifiable via backtracking, tool-use failures frequently induce…

Forecasting is a challenging task that offers a clearly measurable way to study AI systems. Forecasting requires a large amount of research on the internet, and evaluations require time for events to happen, making the development of…

Numerous software analysis tools exist today, yet applying them to diverse open-source projects remains challenging due to environment setup, dependency resolution, and tool configuration. LLM-based agents offer a potential solution, yet no…

软件工程 · 计算机科学 2026-04-20 Islem Bouzenia , Cristian Cadar , Michael Pradel

Existing benchmarks for tool-using LLM agents primarily report single-run success rates and miss reliability properties required in production. We introduce \textbf{ReliabilityBench}, a benchmark for evaluating agent reliability across…

人工智能 · 计算机科学 2026-01-13 Aayush Gupta

Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployments often that observed on benchmark settings, especially in…

With the advancement of automated software engineering, research focus is increasingly shifting toward practical tasks reflecting the day-to-day work of software engineers. Among these tasks, software migration, a critical process of…

软件工程 · 计算机科学 2026-04-29 Ryo Fujii , Makoto Morishita , Kazuki Yano , Jun Suzuki

Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. However, existing benchmarks remain focused on single agentic capability, failing to capture…

Reward models (RMs) are at the crux of successfully using RLHF to align pretrained models to human preferences, yet there has been relatively little study that focuses on evaluation of those models. Evaluating reward models presents an…