Benchmark AI_TRUECOST_

real AI cost, not vanity metrics

Follow on X updates & details

What does AI actually cost to ship a ticket?

Every harness is tested against every model on the same tickets, so the real cost across all of them is visible.

The problem
Benchmark leaderboards are vanity metrics. They don't show what a model costs, or how it performs inside a real coding harness.
What you get
The exact cost of every harness against every model, run on the same repos and the same tickets.
Why it matters
Save money and time, and pick the best harness for your favorite model.

Select a ticket to load the harness × model cost grid.

Workflow page →

Swipe to view more

Average cost of passing runs for each harness and model on this ticket. The cheapest passing cell is highlighted. Click a price for the run breakdown.
Model codex cursor
composer-2.5 $0.30
gpt-5.6-luna $0.07 $0.10
grok-4.6 $1.28