Skip to content

SWE-Lancer what it measures, how it is scored and who leads it

Coding · OpenAI · introduced Feb 17, 2025 · Dollars earned

About this benchmark

Task
Complete real freelance software jobs from Upwork, both implementation and choosing between proposals.
Dataset
Over 1,400 tasks worth $1 million in real payouts, from $50 bug fixes to $32,000 features.
Method
Implementation tasks graded by end-to-end tests written by engineers; managerial tasks against the original manager's choice.
Metric
Dollars earned
Organization
OpenAI
Introduced
Feb 17, 2025

Versions

Newest first. New versions are added, never rewritten.

VersionDate
1Released with the paper.Feb 17, 20251 year ago

Sources: each benchmark's paper (arXiv) and its official site or leaderboard. Scores are the boards captured here every day, exactly as published.

New #1s by email

Saturdays, only in weeks when a leaderboard has a new #1.

Double opt-in. Unsubscribe any time.