Model Leaderboard

Which DeepSWE model gives you the most performance per dollar — and where you are overpaying for the same score.

Best Models

The most cost-efficient model for each performance score, from all available models.

V4 rates effective Aug 16, 2026
  1. claude-opus-5high· Anthropic
    72.8%±1.9
    $0.083/pt
  2. gpt-5-6-solhigh· 43% cheaper · 3.4% worse than claude-opus-5 / high
    69.4%±1.4
    $0.050/pt
  3. gpt-5-6-lunamax· 13% cheaper · 2.2% worse than gpt-5-6-sol / high
    67.2%±4.0
    $0.045/pt
  4. gemini-3-7-flashmedium· 33% cheaper · 1.7% worse than gpt-5-6-luna / max
    65.5%±3.1
    $0.031/pt
  5. deepseek-v4-promax· 38% cheaper · 2.7% worse than gemini-3-7-flash / medium
    62.8%±6.3
    $0.020/pt
  6. deepseek-v4-flashmax· 72% cheaper · 9.5% worse than deepseek-v4-pro / max
    53.3%±3.6
    0.657¢/pt
PerformanceCost ($1 blocks, 10¢ blocks)UncertainEmpty

Models in the same performance bracket are treated as similar-score alternatives in the comparison.

Charts

Compare deepswe.datacurve.ai pass rate with mean cost. The top-right corner is the sweet spot: higher performance for less money.

15 selected models
Providers

Chart loads in the browser.

Comparison Table

15 models · sort by performance, price, or tokens/turns

Sort comparison table
V4 rates effective Aug 16, 2026
  • claude-opus-5max· Anthropic
    73.6%±3.9
    $0.161/pt
  • gpt-5-6-solmax· 29% cheaper · 1.0% worse than claude-opus-5 / max
    72.7%±2.8
    $0.115/pt
  • claude-fable-5xhigh· 60% pricier · 2.8% worse than gpt-5-6-sol / max
    69.9%±3.2
    $0.192/pt
  • gpt-5-6-terramax· 63% cheaper · 0.3% worse than claude-fable-5 / xhigh
    69.6%±2.6
    $0.071/pt
  • glm-5-3max· 19% cheaper · 0.7% worse than gpt-5-6-terra / max
    69.0%±3.0
    $0.058/pt
  • kimi-k3max· 17% pricier · 0.4% worse than glm-5-3 / max
    68.5%±4.5
    $0.068/pt
  • grok-4-6medium· 26% cheaper · 1.0% worse than kimi-k3 / max
    67.5%±2.3
    $0.051/pt
  • gpt-5-6-lunamax· 12% cheaper · 0.3% worse than grok-4-6 / medium
    67.2%±4.0
    $0.045/pt
  • gemini-3-7-flashmedium· 33% cheaper · 1.7% worse than gpt-5-6-luna / max
    65.5%±3.1
    $0.031/pt
  • deepseek-v4-promax· 38% cheaper · 2.7% worse than gemini-3-7-flash / medium
    62.8%±6.3
    $0.020/pt
  • qwen3-8-maxxhigh· 198% pricier · 5.4% worse than deepseek-v4-pro / max
    57.5%±2.7
    $0.065/pt
  • muse-spark-1-2xhigh· 1% cheaper · 2.6% worse than qwen3-8-max / xhigh
    54.9%±2.1
    $0.067/pt
  • claude-sonnet-5max· 614% pricier · 1.0% worse than muse-spark-1-2 / xhigh
    53.8%±4.2
    $0.490/pt
  • deepseek-v4-flashmax· 99% cheaper · 0.5% worse than claude-sonnet-5 / max
    53.3%±3.6
    0.657¢/pt
  • kimi-k2-7-codedefault· 704% pricier · 22.8% worse than deepseek-v4-flash / max
    30.5%±0.5
    $0.092/pt
PerformanceCost ($1 blocks, 10¢ blocks)UncertainEmpty

Models in the same performance bracket are treated as similar-score alternatives in the comparison.

Model selector

Choose the model and reasoning effort rows to compare.

15 of 44 selected
ProvidersAll providers shown

Alibaba

qwen3-8-max

Alibaba

Anthropic

claude-opus-5

Anthropic

claude-fable-5

Anthropic

claude-sonnet-5

Anthropic

DeepSeek

deepseek-v4-pro

DeepSeek

deepseek-v4-flash

DeepSeek

Google

gemini-3-7-flash

Google

Meta

muse-spark-1-2

Meta

Moonshot

kimi-k3

Moonshot

kimi-k2-7-code

Moonshot

OpenAI

gpt-5-6-sol

OpenAI

gpt-5-6-terra

OpenAI

gpt-5-6-luna

OpenAI

xAI

grok-4-6

xAI

Zhipu

glm-5-3

Zhipu