Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale
Abstract
This study explores Large Language Models (LLMs) as autonomous agents for real-world tasks, including freelance software development. This work presents a new benchmark that evaluates LLMs on freelance programming and data analysis tasks derived from economic data. We construct the benchmark using synthetic tasks created from a Kaggle Freelancer dataset of job postings, with all job prices standardized to USD (median fixed-project price around 306). Each task is accompanied by structured input-output test cases and an estimated price tag, enabling automated correctness checking and a monetary performance valuation. This approach is inspired by OpenAI's recent SWE-Lancer benchmark (1,400 real Upwork tasks worth 1.52 million USD, followed closely by GPT-4o-mini at 1.33M) and Mistral ($0.70M). We analyze the distribution of errors per task and observe that the strongest models solve the most tasks and rarely fail completely on any project. We discuss the implications of these results for the feasibility of AI as a freelance developer, the advantages and limitations of our automated benchmark approach, and the gap between performance on structured tasks versus the true complexity of real-world freelance jobs.
Keywords
Cite
@article{arxiv.2505.13511,
title = {Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale},
author = {David Noever and Forrest McKee},
journal= {arXiv preprint arXiv:2505.13511},
year = {2025}
}