Model Leaderboard

Which DeepSWE model gives you the most performance per dollar — and where you are overpaying for the same score.

Best Models

The most cost-efficient model for each performance score, from all available models.

V4 rates effective Aug 16, 2026
  1. gemini-3-8-flashhigh· Google
    73.8%±1.4
    $0.032/pt
  2. gemini-3-8-flashmedium· 17% cheaper · 2.8% worse than gemini-3-8-flash / high
    71.0%±2.3
    $0.028/pt
  3. GPT-6 Lunamaxsource· 89% cheaper · 4.4% worse than gemini-3-8-flash / medium
    66.6%
    0.326¢/pt
  4. GPT-6 Lunaxhighsource· 49% cheaper · 5.3% worse than GPT-6 Luna / max
    61.3%
    0.179¢/pt
  5. GPT-6 Lunahighsource· 24% cheaper · 2.0% worse than GPT-6 Luna / xhigh
    59.3%
    0.141¢/pt
  6. GPT-6 Lunamediumsource· 38% cheaper · 14.8% worse than GPT-6 Luna / high
    44.5%
    0.116¢/pt
  7. GPT-6 Lunalowsource· 89% cheaper · 42.0% worse than GPT-6 Luna / medium
    2.4%
    0.235¢/pt
PerformanceCost ($1 blocks, 10¢ blocks)UncertainEmpty

Models in the same performance bracket are treated as similar-score alternatives in the comparison.

Charts

Compare deepswe.datacurve.ai pass rate with mean cost. The top-right corner is the sweet spot: higher performance for less money.

19 selected models
Providers

Chart loads in the browser.

Comparison Table

19 models · sort by performance, price, or tokens/turns

Sort comparison table
V4 rates effective Aug 16, 2026
  • gpt-6-astraxhigh· OpenAI
    74.1%±2.9
    $0.088/pt
  • gemini-3-8-flashhigh· 64% cheaper · 0.3% worse than gpt-6-astra / xhigh
    73.8%±1.4
    $0.032/pt
  • claude-opus-5max· 401% pricier · 0.2% worse than gemini-3-8-flash / high
    73.6%±3.9
    $0.161/pt
  • gpt-5-6-solmax· 29% cheaper · 1.0% worse than claude-opus-5 / max
    72.7%±2.8
    $0.115/pt
  • claude-fable-5xhigh· 60% pricier · 2.8% worse than gpt-5-6-sol / max
    69.9%±3.2
    $0.192/pt
  • gpt-5-6-terramax· 63% cheaper · 0.3% worse than claude-fable-5 / xhigh
    69.6%±2.6
    $0.071/pt
  • glm-5-3max· 19% cheaper · 0.7% worse than gpt-5-6-terra / max
    69.0%±3.0
    $0.058/pt
  • GPT-6 Solmaxsource· 31% cheaper · 0.1% worse than glm-5-3 / max
    68.8%
    $0.040/pt
  • kimi-k3max· 70% pricier · 0.3% worse than GPT-6 Sol / max
    68.5%±4.5
    $0.068/pt
  • grok-4-6medium· 26% cheaper · 1.0% worse than kimi-k3 / max
    67.5%±2.3
    $0.051/pt
  • gpt-5-6-lunamax· 12% cheaper · 0.3% worse than grok-4-6 / medium
    67.2%±4.0
    $0.045/pt
  • GPT-6 Lunamaxsource· 93% cheaper · 0.6% worse than gpt-5-6-luna / max
    66.6%
    0.326¢/pt
  • glm-5-3-flashmax· 122% pricier · 3.2% worse than GPT-6 Luna / max
    63.4%±4.4
    0.760¢/pt
  • deepseek-v4-promax· 159% pricier · 0.6% worse than glm-5-3-flash / max
    62.8%±6.3
    $0.020/pt
  • qwen3-8-maxxhigh· 198% pricier · 5.4% worse than deepseek-v4-pro / max
    57.5%±2.7
    $0.065/pt
  • muse-spark-1-2xhigh· 1% cheaper · 2.6% worse than qwen3-8-max / xhigh
    54.9%±2.1
    $0.067/pt
  • claude-sonnet-5max· 614% pricier · 1.0% worse than muse-spark-1-2 / xhigh
    53.8%±4.2
    $0.490/pt
  • deepseek-v4-flashmax· 99% cheaper · 0.5% worse than claude-sonnet-5 / max
    53.3%±3.6
    0.657¢/pt
  • kimi-k2-7-codedefault· 704% pricier · 22.8% worse than deepseek-v4-flash / max
    30.5%±0.5
    $0.092/pt
PerformanceCost ($1 blocks, 10¢ blocks)UncertainEmpty

Models in the same performance bracket are treated as similar-score alternatives in the comparison.

Model selector

Choose the model and reasoning effort rows to compare.

19 of 59 selected
ProvidersAll providers shown

Alibaba

qwen3-8-max

Alibaba

Anthropic

claude-opus-5

Anthropic

claude-fable-5

Anthropic

claude-sonnet-5

Anthropic

DeepSeek

deepseek-v4-pro

DeepSeek

deepseek-v4-flash

DeepSeek

Google

gemini-3-8-flash

Google

Meta

muse-spark-1-2

Meta

Moonshot

kimi-k3

Moonshot

kimi-k2-7-code

Moonshot

OpenAI

gpt-6-astra

OpenAI

gpt-5-6-sol

OpenAI

gpt-5-6-terra

OpenAI

GPT-6 Sol

OpenAI

gpt-5-6-luna

OpenAI

GPT-6 Luna

OpenAI

xAI

grok-4-6

xAI

Zhipu

glm-5-3

Zhipu

glm-5-3-flash

Zhipu