What will your workload actually cost?
Real prices, public benchmarks. Pick a task, set your volume, and see what each hosted model costs to run — plus quality data where public benchmarks exist.
Benchmark data updated July 27, 2026 · Methodology & sources
Quality vs. Cost — Coding / Frontend & UI
Scores from public benchmarks · cost is an EstimateSWE-V-adjusted view: 3 of 15 models shown. Models without SWE-bench Verified data are excluded from this chart, not ranked lower.
Excluded from this view: Claude Opus 4.8, Claude Sonnet 5, Gemini 3.5 Flash and GLM-5.2 +8 more — no eligible SWE-V score.
Not shown on the adjusted axis (no SWE-bench Verified score — we never estimate one): Gemini 3.5 Flash, GLM-5.2, DeepSeek V4 Pro and Kimi K3
Not rated (no public benchmark data for this category): Claude Opus 4.8, Claude Sonnet 5, DeepSeek V4 Flash, Gemini 2.5 Flash, GPT-5.6 Sol, Claude Sonnet 4.5, Claude Haiku 4.5 and Claude Opus 4.5
All Models at a Glance
Prices are our marketplace rates · cost/job is an Estimate| GLM-4.5-Air | Tier 2 · Balanced | $0.20 | $1.10 | 125K | $0.0025est. | $0.0043est.self-reported | 1,372.9 |
| GLM-4.6 | Tier 2 · Balanced | $0.60 | $2.20 | 200K | $0.0057est. | $0.0084est. | 1,425.2 |
| GLM-5 | Tier 2 · Balanced | $1.00 | $3.20 | 200K | $0.0088est. | $0.01est. | 1,456.7 |
| Claude Opus 4.8 | — | $5.00 | $25.00 | 200K | $0.06est. | No SWE-V data | — |
| Claude Sonnet 5 | — | $3.00 | $15.00 | 200K | $0.03est. | No SWE-V data | — |
| Gemini 3.5 Flash | Tier 2 · Balanced | $1.50 | $9.00 | 1M | $0.02est. | No SWE-V data | 1,474.3 |
| GLM-5.2 | Tier 2 · Balanced | $1.40 | $4.40 | 200K | $0.01est. | No SWE-V data | 1,469.2 |
| DeepSeek V4 Flash | — | $0.14 | $0.28 | 128K | $0.0010est. | No SWE-V data | 1,435.5 |
| DeepSeek V4 Pro | Tier 2 · Balanced | $0.44 | $0.87 | 128K | $0.0030est. | No SWE-V data | 1,456.6 |
| Gemini 2.5 Flash | — | $0.30 | $2.50 | 1M | $0.0049est. | No SWE-V data | 1,410.2 |
| GPT-5.6 Sol | — | $5.00 | $30.00 | 400K | $0.07est. | No SWE-V data | — |
| Kimi K3 | Tier 1 · Frontier | $3.00 | $15.00 | 1M | $0.03est. | No SWE-V data | 1,485.7 |
| Claude Sonnet 4.5 | — | $3.00 | $15.00 | 200K | $0.03est. | No SWE-V data | — |
| Claude Haiku 4.5 | — | $0.80 | $4.00 | 200K | $0.0092est. | No SWE-V data | — |
| Claude Opus 4.5 | — | $15.00 | $75.00 | 200K | $0.17est. | No SWE-V data | — |
Cost-per-Job Calculator
Single code generation / edit task
SWE-V-adjusted view: 3 of 15 models shown. Models without SWE-bench Verified data are excluded from this chart, not ranked lower.
Excluded from this view: Claude Opus 4.8, Claude Sonnet 5, Gemini 3.5 Flash and GLM-5.2 +8 more — no eligible SWE-V score.
Not shown on the SWE-V-adjusted basis (no SWE-bench Verified score — we never estimate one): Claude Opus 4.8, Claude Sonnet 5, Gemini 3.5 Flash, GLM-5.2, DeepSeek V4 Flash, DeepSeek V4 Pro, Gemini 2.5 Flash, GPT-5.6 Sol, Kimi K3, Claude Sonnet 4.5, Claude Haiku 4.5 and Claude Opus 4.5
Methodology & Data Sources
Data last changed July 27, 2026; each score carries its own as-of date below (no-change sync runs do not bump this date, and upstream leaderboards publish on their own schedules). Prices are our marketplace rates, fetched live. Benchmark scores come from public publications and leaderboards — never our own measurements, and never fabricated: where no direct public score exists, we show nothing.
Benchmarks used
| Benchmark | Metric | Anchor low → high | Source |
|---|---|---|---|
| SWE-bench Verified | pass_rate_pct | 40 → 85 | https://www.swebench.com/ |
| LMArena WebDev | elo | 1,200 → 1,700 | https://arena.ai/leaderboard |
| LiveBench Coding | score_0_100 | 40 → 85 | https://livebench.ai/ |
Quality scores normalize each raw benchmark value with fixed anchors — norm = clamp((raw − anchor_low) / (anchor_high − anchor_low), 0, 1) × 100 — then take the category’s weighted average. When a model is missing a benchmark, the remaining weights are re-normalized; a model with no scored benchmark in a category gets no quality score at all.
Every score we use, with its source
- DeepSeek V4 Flash — LiveBench Coding: 69.228 source (as of June 25, 2026)LiveBench Coding category average (code_generation + code_completion), snapshot 2026-06-25.
- DeepSeek V4 Pro — LiveBench Coding: 69.994 source (as of June 25, 2026)LiveBench Coding category average (code_generation + code_completion), snapshot 2026-06-25.
- Gemini 3.5 Flash — LiveBench Coding: 78.185 source (as of June 25, 2026)LiveBench Coding category average (code_generation + code_completion), snapshot 2026-06-25. LiveBench entry is gemini-3.5-flash-high (reasoning-effort "high" mode).
- GLM-5.2 — LiveBench Coding: 79.654 source (as of June 25, 2026)LiveBench Coding category average (code_generation + code_completion), snapshot 2026-06-25.
- Kimi K2.6 — LiveBench Coding: 78.567 source (as of June 25, 2026)LiveBench Coding category average (code_generation + code_completion), snapshot 2026-06-25. LiveBench entry is the "thinking" mode variant.
- DeepSeek V4 Flash — LMArena Text (overall Elo): 1,435.5 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
- DeepSeek V4 Pro — LMArena Text (overall Elo): 1,456.6 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
- Gemini 2.5 Flash — LMArena Text (overall Elo): 1,410.2 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
- Gemini 3.5 Flash — LMArena Text (overall Elo): 1,474.3 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall". Arena entry is gemini-3.5-flash-medium (reasoning-effort "medium" mode).
- GLM-4.5-Air — LMArena Text (overall Elo): 1,372.9 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
- GLM-4.6 — LMArena Text (overall Elo): 1,425.2 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
- GLM-5 — LMArena Text (overall Elo): 1,456.7 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
- GLM-5.2 — LMArena Text (overall Elo): 1,469.2 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall". Arena entry is the "(max)" mode variant.
- Kimi K2.6 — LMArena Text (overall Elo): 1,460.9 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
- Kimi K3 — LMArena Text (overall Elo): 1,485.7 source (as of July 21, 2026)LMArena Text leaderboard (style control), category "overall".
- DeepSeek V4 Pro — LMArena WebDev: 1,446.7 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall".
- Gemini 3.5 Flash — LMArena WebDev: 1,484.3 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall". Arena entry is gemini-3.5-flash-medium (reasoning-effort "medium" mode).
- GLM-4.6 — LMArena WebDev: 1,338.7 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall".
- GLM-5 — LMArena WebDev: 1,434.7 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall".
- GLM-5.2 — LMArena WebDev: 1,587.7 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall". Arena entry is the "(max)" mode variant.
- Kimi K2.6 — LMArena WebDev: 1,509.8 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall".
- Kimi K3 — LMArena WebDev: 1,682 source (as of July 24, 2026)LMArena WebDev leaderboard, category "overall".
- GLM-4.5-Air — SWE-bench Verified: 57.6 source (as of July 23, 2026)self-reportedGLM-4.5 technical report (Air variant). Vendor self-reported; agent harnesses differ across vendors, so scores are not strictly comparable.
- GLM-4.6 — SWE-bench Verified: 68.2 source (as of September 30, 2025)SWE-bench Verified leaderboard. Submission logs not yet checked by SWE-bench maintainers. Official SWE-bench Verified leaderboard submission (vendor scaffold).
- GLM-5 — SWE-bench Verified: 72.8 source (as of February 17, 2026)SWE-bench Verified leaderboard. No check status is published for this submission in the fetched leaderboard data. mini-SWE-agent harness, high reasoning effort.
- Kimi K2.6 — SWE-bench Verified: 80.2 source (as of July 23, 2026)self-reportedOfficial Moonshot forum announcement. Vendor self-reported; agent harnesses differ across vendors, so scores are not strictly comparable.
SWE-bench scores mix official-leaderboard entries with vendor self-reported results (flagged per score above); agent harnesses differ, so scores are not strictly comparable across models.
SWE-V-adjusted cost per job
For models with a SWE-bench Verified score we also show swev_adjusted_cost = cost_per_job ÷ (SWE-bench Verified pass rate / 100) — the estimated cost per solved task if failures were blindly retried. Only the SWE-bench Verified pass rate is ever used as the denominator; Arena WebDev and LiveBench scores have no probability semantics and are never substituted. Models without a SWE-V score show “—” — we never estimate one.
- The 1/p formula assumes independent retries at a constant success rate.
- It assumes each retry uses the same tokens as the first attempt.
- Real failures often need human or tool intervention rather than a blind retry — that overhead is not priced in.
- Solve rates come from the SWE-bench Verified task set, which differs from your task distribution.
Capability tiers
Models with a quality score are grouped into 1-3 tiers (Tier 1 Frontier / Tier 2 Balanced / Tier 3 Budget) by a deterministic natural-gap rule: scores are rounded to 1 decimal, then any gap between adjacent models (in descending order) of at least max(score range × 25%, 8 points) is a candidate tier boundary; only the largest two candidate gaps become boundaries (ties go to the higher-ranked position). No qualifying gap means a single tier. Tier boundaries are pinned by a build-time snapshot test, so a data update can never silently move a model between tiers — every tier change is human-reviewed. The “Cheapest in Tier 1” highlight marks the lowest cost-per-job model inside the top tier; tier membership only requires a quality score, so a model can hold a tier while its SWE-V-adjusted cost column shows “—”.
Why most categories have no quality ranking
We only publish a quality ranking for a category when at least 5 of our models have a direct public score — exact model version, exact benchmark version, eval mode noted. Right now that bar is met only for Coding / Frontend & UI. For every other category the public leaderboards either don’t cover our catalog or have stopped updating, so we show the cost calculator only. We would rather show you nothing than a made-up number.
Task profiles behind the calculator
| Category | Input / output tokens | Assumption |
|---|---|---|
| Coding / Frontend & UI | 4,000 / 1,500 | Read a medium file of surrounding context and produce a function-level patch. |
| Writing & Marketing Copy | 800 / 1,200 | Brief plus brand notes in, one short marketing text or email out. |
| Data Extraction & JSON Output | 2,500 / 400 | A page of unstructured text in, fixed-schema JSON out. |
| Agent Tool Calling | 6,000 / 800 | Multi-turn tool-calling session, including tool schemas and intermediate results fed back into context. |
| Long-Document RAG / Summarization | 30,000 / 1,500 | Tens of pages of documents in, one structured summary out. |
| Translation | 1,500 / 1,600 | Chinese/English translation of a ~1000-word article; output tokens roughly match input. |
Estimates: cost-per-job figures are estimates based on the typical token profiles above — actual usage varies. Prices are our real marketplace rates.
Arena Text Elo column: general chat Elo from LMArena (community votes), not task-specific; shown for reference only and never used in quality scores or badges.
Freshness: benchmark scores are reviewed monthly (each carries its own as-of date); prices reflect our current marketplace rates.
Independence: scores come from public publications, not our own testing. Vendor self-reported numbers are labeled as such wherever they appear.