What will your workload actually cost?

Real prices, public benchmarks. Pick a task, set your volume, and see what each hosted model costs to run — plus quality data where public benchmarks exist.

Benchmark data updated July 27, 2026 · Methodology & sources

All Models at a Glance

Prices are our marketplace rates · cost/job is an Estimate
Kimi K2.6Tier 1 · Frontier$0.76$3.20256K$0.0078est.$0.0098est.self-reported1,460.9
GLM-5Tier 2 · Balanced$1.00$3.20200K$0.0088est.$0.01est.1,456.7
DeepSeek V4 Flash$0.20$0.591M$0.0017est.No SWE-V data1,435.5
DeepSeek V4 ProTier 3 · Budget$0.59$1.781M$0.0050est.No SWE-V data1,456.6
GLM-5-Turbo$0.96$3.20200K$0.0086est.No SWE-V data
Qwen3.7-Plus$0.32$1.281M$0.0032est.No SWE-V data
Qwen3.7-Max$2.00$6.001M$0.02est.No SWE-V data
MiniMax M3$0.24$0.961M$0.0024est.No SWE-V data

Cost-per-Job Calculator

Single code generation / edit task

Estimate — based on typical token profiles, actual usage varies
Cost basis:

SWE-V-adjusted view: 2 of 8 models shown. Models without SWE-bench Verified data are excluded from this chart, not ranked lower.

Excluded from this view: DeepSeek V4 Flash, DeepSeek V4 Pro, GLM-5-Turbo and Qwen3.7-Plus +2 more — no eligible SWE-V score.

1,000
101001k10k100k1M
Cheapest per solved taskCheapest in Tier 1 — cheapest model in the top capability tier

Not shown on the SWE-V-adjusted basis (no SWE-bench Verified score — we never estimate one): DeepSeek V4 Flash, DeepSeek V4 Pro, GLM-5-Turbo, Qwen3.7-Plus, Qwen3.7-Max and MiniMax M3

Methodology & Data Sources

Data last changed July 27, 2026; each score carries its own as-of date below (no-change sync runs do not bump this date, and upstream leaderboards publish on their own schedules). Prices are our marketplace rates, fetched live. Benchmark scores come from public publications and leaderboards — never our own measurements, and never fabricated: where no direct public score exists, we show nothing.

Benchmarks used

BenchmarkMetricAnchor low → highSource
SWE-bench Verifiedpass_rate_pct4085https://www.swebench.com/
LMArena WebDevelo1,2001,700https://arena.ai/leaderboard
LiveBench Codingscore_0_1004085https://livebench.ai/

Quality scores normalize each raw benchmark value with fixed anchors — norm = clamp((raw − anchor_low) / (anchor_high − anchor_low), 0, 1) × 100 — then take the category’s weighted average. When a model is missing a benchmark, the remaining weights are re-normalized; a model with no scored benchmark in a category gets no quality score at all.

SWE-V-adjusted cost per job

For models with a SWE-bench Verified score we also show swev_adjusted_cost = cost_per_job ÷ (SWE-bench Verified pass rate / 100) — the estimated cost per solved task if failures were blindly retried. Only the SWE-bench Verified pass rate is ever used as the denominator; Arena WebDev and LiveBench scores have no probability semantics and are never substituted. Models without a SWE-V score show “—” — we never estimate one.

  • The 1/p formula assumes independent retries at a constant success rate.
  • It assumes each retry uses the same tokens as the first attempt.
  • Real failures often need human or tool intervention rather than a blind retry — that overhead is not priced in.
  • Solve rates come from the SWE-bench Verified task set, which differs from your task distribution.

Capability tiers

Models with a quality score are grouped into 1-3 tiers (Tier 1 Frontier / Tier 2 Balanced / Tier 3 Budget) by a deterministic natural-gap rule: scores are rounded to 1 decimal, then any gap between adjacent models (in descending order) of at least max(score range × 25%, 8 points) is a candidate tier boundary; only the largest two candidate gaps become boundaries (ties go to the higher-ranked position). No qualifying gap means a single tier. Tier boundaries are pinned by a build-time snapshot test, so a data update can never silently move a model between tiers — every tier change is human-reviewed. The “Cheapest in Tier 1” highlight marks the lowest cost-per-job model inside the top tier; tier membership only requires a quality score, so a model can hold a tier while its SWE-V-adjusted cost column shows “—”.

Why most categories have no quality ranking

We only publish a quality ranking for a category when at least 5 of our models have a direct public score — exact model version, exact benchmark version, eval mode noted. Right now that bar is met only for Coding / Frontend & UI. For every other category the public leaderboards either don’t cover our catalog or have stopped updating, so we show the cost calculator only. We would rather show you nothing than a made-up number.

Task profiles behind the calculator

CategoryInput / output tokensAssumption
Coding / Frontend & UI4,000 / 1,500Read a medium file of surrounding context and produce a function-level patch.
Writing & Marketing Copy800 / 1,200Brief plus brand notes in, one short marketing text or email out.
Data Extraction & JSON Output2,500 / 400A page of unstructured text in, fixed-schema JSON out.
Agent Tool Calling6,000 / 800Multi-turn tool-calling session, including tool schemas and intermediate results fed back into context.
Long-Document RAG / Summarization30,000 / 1,500Tens of pages of documents in, one structured summary out.
Translation1,500 / 1,600Chinese/English translation of a ~1000-word article; output tokens roughly match input.

Estimates: cost-per-job figures are estimates based on the typical token profiles above — actual usage varies. Prices are our real marketplace rates.

Arena Text Elo column: general chat Elo from LMArena (community votes), not task-specific; shown for reference only and never used in quality scores or badges.

Freshness: benchmark scores are reviewed monthly (each carries its own as-of date); prices reflect our current marketplace rates.

Independence: scores come from public publications, not our own testing. Vendor self-reported numbers are labeled as such wherever they appear.