What will your workload actually cost?
Real prices, public benchmarks. Pick a task, set your volume, and see what each hosted model costs to run — plus quality data where public benchmarks exist.
Benchmark data updated July 27, 2026 · Methodology & sources
All Models at a Glance
Prices are our marketplace rates · cost/job is an Estimate| Kimi K2.6 | Tier 1 · Frontier | $0.76 | $3.20 | 256K | $0.0078est. | $0.0098est.self-reported | 1,460.9 |
| GLM-5 | Tier 2 · Balanced | $1.00 | $3.20 | 200K | $0.0088est. | $0.01est. | 1,456.7 |
| DeepSeek V4 Flash | — | $0.20 | $0.59 | 1M | $0.0017est. | No SWE-V data | 1,435.5 |
| DeepSeek V4 Pro | Tier 3 · Budget | $0.59 | $1.78 | 1M | $0.0050est. | No SWE-V data | 1,456.6 |
| GLM-5-Turbo | — | $0.96 | $3.20 | 200K | $0.0086est. | No SWE-V data | — |
| Qwen3.7-Plus | — | $0.32 | $1.28 | 1M | $0.0032est. | No SWE-V data | — |
| Qwen3.7-Max | — | $2.00 | $6.00 | 1M | $0.02est. | No SWE-V data | — |
| MiniMax M3 | — | $0.24 | $0.96 | 1M | $0.0024est. | No SWE-V data | — |
Cost-per-Job Calculator
Single code generation / edit task
SWE-V-adjusted view: 2 of 8 models shown. Models without SWE-bench Verified data are excluded from this chart, not ranked lower.
Excluded from this view: DeepSeek V4 Flash, DeepSeek V4 Pro, GLM-5-Turbo and Qwen3.7-Plus +2 more — no eligible SWE-V score.
Not shown on the SWE-V-adjusted basis (no SWE-bench Verified score — we never estimate one): DeepSeek V4 Flash, DeepSeek V4 Pro, GLM-5-Turbo, Qwen3.7-Plus, Qwen3.7-Max and MiniMax M3
Methodology & Data Sources
Data last changed July 27, 2026; each score carries its own as-of date below (no-change sync runs do not bump this date, and upstream leaderboards publish on their own schedules). Prices are our marketplace rates, fetched live. Benchmark scores come from public publications and leaderboards — never our own measurements, and never fabricated: where no direct public score exists, we show nothing.
Benchmarks used
| Benchmark | Metric | Anchor low → high | Source |
|---|---|---|---|
| SWE-bench Verified | pass_rate_pct | 40 → 85 | https://www.swebench.com/ |
| LMArena WebDev | elo | 1,200 → 1,700 | https://arena.ai/leaderboard |
| LiveBench Coding | score_0_100 | 40 → 85 | https://livebench.ai/ |
Quality scores normalize each raw benchmark value with fixed anchors — norm = clamp((raw − anchor_low) / (anchor_high − anchor_low), 0, 1) × 100 — then take the category’s weighted average. When a model is missing a benchmark, the remaining weights are re-normalized; a model with no scored benchmark in a category gets no quality score at all.
SWE-V-adjusted cost per job
For models with a SWE-bench Verified score we also show swev_adjusted_cost = cost_per_job ÷ (SWE-bench Verified pass rate / 100) — the estimated cost per solved task if failures were blindly retried. Only the SWE-bench Verified pass rate is ever used as the denominator; Arena WebDev and LiveBench scores have no probability semantics and are never substituted. Models without a SWE-V score show “—” — we never estimate one.
- The 1/p formula assumes independent retries at a constant success rate.
- It assumes each retry uses the same tokens as the first attempt.
- Real failures often need human or tool intervention rather than a blind retry — that overhead is not priced in.
- Solve rates come from the SWE-bench Verified task set, which differs from your task distribution.
Capability tiers
Models with a quality score are grouped into 1-3 tiers (Tier 1 Frontier / Tier 2 Balanced / Tier 3 Budget) by a deterministic natural-gap rule: scores are rounded to 1 decimal, then any gap between adjacent models (in descending order) of at least max(score range × 25%, 8 points) is a candidate tier boundary; only the largest two candidate gaps become boundaries (ties go to the higher-ranked position). No qualifying gap means a single tier. Tier boundaries are pinned by a build-time snapshot test, so a data update can never silently move a model between tiers — every tier change is human-reviewed. The “Cheapest in Tier 1” highlight marks the lowest cost-per-job model inside the top tier; tier membership only requires a quality score, so a model can hold a tier while its SWE-V-adjusted cost column shows “—”.
Why most categories have no quality ranking
We only publish a quality ranking for a category when at least 5 of our models have a direct public score — exact model version, exact benchmark version, eval mode noted. Right now that bar is met only for Coding / Frontend & UI. For every other category the public leaderboards either don’t cover our catalog or have stopped updating, so we show the cost calculator only. We would rather show you nothing than a made-up number.
Task profiles behind the calculator
| Category | Input / output tokens | Assumption |
|---|---|---|
| Coding / Frontend & UI | 4,000 / 1,500 | Read a medium file of surrounding context and produce a function-level patch. |
| Writing & Marketing Copy | 800 / 1,200 | Brief plus brand notes in, one short marketing text or email out. |
| Data Extraction & JSON Output | 2,500 / 400 | A page of unstructured text in, fixed-schema JSON out. |
| Agent Tool Calling | 6,000 / 800 | Multi-turn tool-calling session, including tool schemas and intermediate results fed back into context. |
| Long-Document RAG / Summarization | 30,000 / 1,500 | Tens of pages of documents in, one structured summary out. |
| Translation | 1,500 / 1,600 | Chinese/English translation of a ~1000-word article; output tokens roughly match input. |
Estimates: cost-per-job figures are estimates based on the typical token profiles above — actual usage varies. Prices are our real marketplace rates.
Arena Text Elo column: general chat Elo from LMArena (community votes), not task-specific; shown for reference only and never used in quality scores or badges.
Freshness: benchmark scores are reviewed monthly (each carries its own as-of date); prices reflect our current marketplace rates.
Independence: scores come from public publications, not our own testing. Vendor self-reported numbers are labeled as such wherever they appear.