What will your workload actually cost?

Real prices, public benchmarks. Pick a task, set your volume, and see what each hosted model costs to run — plus quality data where public benchmarks exist.

Benchmark data updated July 27, 2026 · Methodology & sources

Quality vs. Cost — Coding / Frontend & UI

Scores from public benchmarks · cost is an Estimate
Cost axis:

SWE-V-adjusted view: 3 of 15 models shown. Models without SWE-bench Verified data are excluded from this chart, not ranked lower.

Excluded from this view: Claude Opus 4.8, Claude Sonnet 5, Gemini 3.5 Flash and GLM-5.2 +8 more — no eligible SWE-V score.

Pareto efficient (no cheaper-and-better alternative)Bubble size = context window

Not shown on the adjusted axis (no SWE-bench Verified score — we never estimate one): Gemini 3.5 Flash, GLM-5.2, DeepSeek V4 Pro and Kimi K3

Not rated (no public benchmark data for this category): Claude Opus 4.8, Claude Sonnet 5, DeepSeek V4 Flash, Gemini 2.5 Flash, GPT-5.6 Sol, Claude Sonnet 4.5, Claude Haiku 4.5 and Claude Opus 4.5

All Models at a Glance

Prices are our marketplace rates · cost/job is an Estimate
GLM-4.5-AirTier 2 · Balanced$0.20$1.10125K$0.0025est.$0.0043est.self-reported1,372.9
GLM-4.6Tier 2 · Balanced$0.60$2.20200K$0.0057est.$0.0084est.1,425.2
GLM-5Tier 2 · Balanced$1.00$3.20200K$0.0088est.$0.01est.1,456.7
Claude Opus 4.8$5.00$25.00200K$0.06est.No SWE-V data
Claude Sonnet 5$3.00$15.00200K$0.03est.No SWE-V data
Gemini 3.5 FlashTier 2 · Balanced$1.50$9.001M$0.02est.No SWE-V data1,474.3
GLM-5.2Tier 2 · Balanced$1.40$4.40200K$0.01est.No SWE-V data1,469.2
DeepSeek V4 Flash$0.14$0.28128K$0.0010est.No SWE-V data1,435.5
DeepSeek V4 ProTier 2 · Balanced$0.44$0.87128K$0.0030est.No SWE-V data1,456.6
Gemini 2.5 Flash$0.30$2.501M$0.0049est.No SWE-V data1,410.2
GPT-5.6 Sol$5.00$30.00400K$0.07est.No SWE-V data
Kimi K3Tier 1 · Frontier$3.00$15.001M$0.03est.No SWE-V data1,485.7
Claude Sonnet 4.5$3.00$15.00200K$0.03est.No SWE-V data
Claude Haiku 4.5$0.80$4.00200K$0.0092est.No SWE-V data
Claude Opus 4.5$15.00$75.00200K$0.17est.No SWE-V data

Cost-per-Job Calculator

Single code generation / edit task

Estimate — based on typical token profiles, actual usage varies
Cost basis:

SWE-V-adjusted view: 3 of 15 models shown. Models without SWE-bench Verified data are excluded from this chart, not ranked lower.

Excluded from this view: Claude Opus 4.8, Claude Sonnet 5, Gemini 3.5 Flash and GLM-5.2 +8 more — no eligible SWE-V score.

1,000
101001k10k100k1M
Cheapest per solved task

Not shown on the SWE-V-adjusted basis (no SWE-bench Verified score — we never estimate one): Claude Opus 4.8, Claude Sonnet 5, Gemini 3.5 Flash, GLM-5.2, DeepSeek V4 Flash, DeepSeek V4 Pro, Gemini 2.5 Flash, GPT-5.6 Sol, Kimi K3, Claude Sonnet 4.5, Claude Haiku 4.5 and Claude Opus 4.5

Methodology & Data Sources

Data last changed July 27, 2026; each score carries its own as-of date below (no-change sync runs do not bump this date, and upstream leaderboards publish on their own schedules). Prices are our marketplace rates, fetched live. Benchmark scores come from public publications and leaderboards — never our own measurements, and never fabricated: where no direct public score exists, we show nothing.

Benchmarks used

BenchmarkMetricAnchor low → highSource
SWE-bench Verifiedpass_rate_pct4085https://www.swebench.com/
LMArena WebDevelo1,2001,700https://arena.ai/leaderboard
LiveBench Codingscore_0_1004085https://livebench.ai/

Quality scores normalize each raw benchmark value with fixed anchors — norm = clamp((raw − anchor_low) / (anchor_high − anchor_low), 0, 1) × 100 — then take the category’s weighted average. When a model is missing a benchmark, the remaining weights are re-normalized; a model with no scored benchmark in a category gets no quality score at all.

Every score we use, with its source

  • DeepSeek V4 FlashLiveBench Coding: 69.228 source (as of June 25, 2026)LiveBench Coding category average (code_generation + code_completion), snapshot 2026-06-25.
  • DeepSeek V4 ProLiveBench Coding: 69.994 source (as of June 25, 2026)LiveBench Coding category average (code_generation + code_completion), snapshot 2026-06-25.
  • Gemini 3.5 FlashLiveBench Coding: 78.185 source (as of June 25, 2026)LiveBench Coding category average (code_generation + code_completion), snapshot 2026-06-25. LiveBench entry is gemini-3.5-flash-high (reasoning-effort "high" mode).
  • GLM-5.2LiveBench Coding: 79.654 source (as of June 25, 2026)LiveBench Coding category average (code_generation + code_completion), snapshot 2026-06-25.
  • Kimi K2.6LiveBench Coding: 78.567 source (as of June 25, 2026)LiveBench Coding category average (code_generation + code_completion), snapshot 2026-06-25. LiveBench entry is the "thinking" mode variant.
  • DeepSeek V4 FlashLMArena Text (overall Elo): 1,435.5 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
  • DeepSeek V4 ProLMArena Text (overall Elo): 1,456.6 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
  • Gemini 2.5 FlashLMArena Text (overall Elo): 1,410.2 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
  • Gemini 3.5 FlashLMArena Text (overall Elo): 1,474.3 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall". Arena entry is gemini-3.5-flash-medium (reasoning-effort "medium" mode).
  • GLM-4.5-AirLMArena Text (overall Elo): 1,372.9 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
  • GLM-4.6LMArena Text (overall Elo): 1,425.2 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
  • GLM-5LMArena Text (overall Elo): 1,456.7 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
  • GLM-5.2LMArena Text (overall Elo): 1,469.2 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall". Arena entry is the "(max)" mode variant.
  • Kimi K2.6LMArena Text (overall Elo): 1,460.9 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
  • Kimi K3LMArena Text (overall Elo): 1,485.7 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
  • DeepSeek V4 ProLMArena WebDev: 1,446.7 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall".
  • Gemini 3.5 FlashLMArena WebDev: 1,484.3 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall". Arena entry is gemini-3.5-flash-medium (reasoning-effort "medium" mode).
  • GLM-4.6LMArena WebDev: 1,338.7 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall".
  • GLM-5LMArena WebDev: 1,434.7 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall".
  • GLM-5.2LMArena WebDev: 1,587.7 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall". Arena entry is the "(max)" mode variant.
  • Kimi K2.6LMArena WebDev: 1,509.8 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall".
  • Kimi K3LMArena WebDev: 1,682 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall".
  • GLM-4.5-AirSWE-bench Verified: 57.6 source (as of July 23, 2026)self-reportedGLM-4.5 technical report (Air variant). Vendor self-reported; agent harnesses differ across vendors, so scores are not strictly comparable.
  • GLM-4.6SWE-bench Verified: 68.2 source (as of September 30, 2025)SWE-bench Verified leaderboard. Submission logs not yet checked by SWE-bench maintainers. Official SWE-bench Verified leaderboard submission (vendor scaffold).
  • GLM-5SWE-bench Verified: 72.8 source (as of February 17, 2026)SWE-bench Verified leaderboard. No check status is published for this submission in the fetched leaderboard data. mini-SWE-agent harness, high reasoning effort.
  • Kimi K2.6SWE-bench Verified: 80.2 source (as of July 23, 2026)self-reportedOfficial Moonshot forum announcement. Vendor self-reported; agent harnesses differ across vendors, so scores are not strictly comparable.

SWE-bench scores mix official-leaderboard entries with vendor self-reported results (flagged per score above); agent harnesses differ, so scores are not strictly comparable across models.

SWE-V-adjusted cost per job

For models with a SWE-bench Verified score we also show swev_adjusted_cost = cost_per_job ÷ (SWE-bench Verified pass rate / 100) — the estimated cost per solved task if failures were blindly retried. Only the SWE-bench Verified pass rate is ever used as the denominator; Arena WebDev and LiveBench scores have no probability semantics and are never substituted. Models without a SWE-V score show “—” — we never estimate one.

  • The 1/p formula assumes independent retries at a constant success rate.
  • It assumes each retry uses the same tokens as the first attempt.
  • Real failures often need human or tool intervention rather than a blind retry — that overhead is not priced in.
  • Solve rates come from the SWE-bench Verified task set, which differs from your task distribution.

Capability tiers

Models with a quality score are grouped into 1-3 tiers (Tier 1 Frontier / Tier 2 Balanced / Tier 3 Budget) by a deterministic natural-gap rule: scores are rounded to 1 decimal, then any gap between adjacent models (in descending order) of at least max(score range × 25%, 8 points) is a candidate tier boundary; only the largest two candidate gaps become boundaries (ties go to the higher-ranked position). No qualifying gap means a single tier. Tier boundaries are pinned by a build-time snapshot test, so a data update can never silently move a model between tiers — every tier change is human-reviewed. The “Cheapest in Tier 1” highlight marks the lowest cost-per-job model inside the top tier; tier membership only requires a quality score, so a model can hold a tier while its SWE-V-adjusted cost column shows “—”.

Why most categories have no quality ranking

We only publish a quality ranking for a category when at least 5 of our models have a direct public score — exact model version, exact benchmark version, eval mode noted. Right now that bar is met only for Coding / Frontend & UI. For every other category the public leaderboards either don’t cover our catalog or have stopped updating, so we show the cost calculator only. We would rather show you nothing than a made-up number.

Task profiles behind the calculator

CategoryInput / output tokensAssumption
Coding / Frontend & UI4,000 / 1,500Read a medium file of surrounding context and produce a function-level patch.
Writing & Marketing Copy800 / 1,200Brief plus brand notes in, one short marketing text or email out.
Data Extraction & JSON Output2,500 / 400A page of unstructured text in, fixed-schema JSON out.
Agent Tool Calling6,000 / 800Multi-turn tool-calling session, including tool schemas and intermediate results fed back into context.
Long-Document RAG / Summarization30,000 / 1,500Tens of pages of documents in, one structured summary out.
Translation1,500 / 1,600Chinese/English translation of a ~1000-word article; output tokens roughly match input.

Estimates: cost-per-job figures are estimates based on the typical token profiles above — actual usage varies. Prices are our real marketplace rates.

Arena Text Elo column: general chat Elo from LMArena (community votes), not task-specific; shown for reference only and never used in quality scores or badges.

Freshness: benchmark scores are reviewed monthly (each carries its own as-of date); prices reflect our current marketplace rates.

Independence: scores come from public publications, not our own testing. Vendor self-reported numbers are labeled as such wherever they appear.