- claude-opus-5high· Anthropic72.8%±1.9$0.083/pt
- gpt-5-6-solhigh· 43% cheaper · 3.4% worse than claude-opus-5 / high69.4%±1.4$0.050/pt
- gpt-5-6-lunamax· 13% cheaper · 2.2% worse than gpt-5-6-sol / high67.2%±4.0$0.045/pt
- gemini-3-7-flashmedium· 33% cheaper · 1.7% worse than gpt-5-6-luna / max65.5%±3.1$0.031/pt
- deepseek-v4-promax· 38% cheaper · 2.7% worse than gemini-3-7-flash / medium62.8%±6.3$0.020/pt
- deepseek-v4-flashmax· 72% cheaper · 9.5% worse than deepseek-v4-pro / max53.3%±3.60.657¢/pt
Model Leaderboard
Which DeepSWE model gives you the most performance per dollar — and where you are overpaying for the same score.
Best Models
The most cost-efficient model for each performance score, from all available models.
V4 rates effective Aug 16, 2026
Charts
Compare deepswe.datacurve.ai pass rate with mean cost. The top-right corner is the sweet spot: higher performance for less money.
15 selected models
Chart loads in the browser.
Comparison Table
15 models · sort by performance, price, or tokens/turns
V4 rates effective Aug 16, 2026
- claude-opus-5max· Anthropic73.6%±3.9$0.161/pt
- gpt-5-6-solmax· 29% cheaper · 1.0% worse than claude-opus-5 / max72.7%±2.8$0.115/pt
- claude-fable-5xhigh· 60% pricier · 2.8% worse than gpt-5-6-sol / max69.9%±3.2$0.192/pt
- gpt-5-6-terramax· 63% cheaper · 0.3% worse than claude-fable-5 / xhigh69.6%±2.6$0.071/pt
- glm-5-3max· 19% cheaper · 0.7% worse than gpt-5-6-terra / max69.0%±3.0$0.058/pt
- kimi-k3max· 17% pricier · 0.4% worse than glm-5-3 / max68.5%±4.5$0.068/pt
- grok-4-6medium· 26% cheaper · 1.0% worse than kimi-k3 / max67.5%±2.3$0.051/pt
- gpt-5-6-lunamax· 12% cheaper · 0.3% worse than grok-4-6 / medium67.2%±4.0$0.045/pt
- gemini-3-7-flashmedium· 33% cheaper · 1.7% worse than gpt-5-6-luna / max65.5%±3.1$0.031/pt
- deepseek-v4-promax· 38% cheaper · 2.7% worse than gemini-3-7-flash / medium62.8%±6.3$0.020/pt
- qwen3-8-maxxhigh· 198% pricier · 5.4% worse than deepseek-v4-pro / max57.5%±2.7$0.065/pt
- muse-spark-1-2xhigh· 1% cheaper · 2.6% worse than qwen3-8-max / xhigh54.9%±2.1$0.067/pt
- claude-sonnet-5max· 614% pricier · 1.0% worse than muse-spark-1-2 / xhigh53.8%±4.2$0.490/pt
- deepseek-v4-flashmax· 99% cheaper · 0.5% worse than claude-sonnet-5 / max53.3%±3.60.657¢/pt
- kimi-k2-7-codedefault· 704% pricier · 22.8% worse than deepseek-v4-flash / max30.5%±0.5$0.092/pt
Model selector
Choose the model and reasoning effort rows to compare.
15 of 44 selected
ProvidersAll providers shown
Alibaba
qwen3-8-max
Alibaba
Anthropic
claude-opus-5
Anthropic
claude-fable-5
Anthropic
claude-sonnet-5
Anthropic
DeepSeek
deepseek-v4-pro
DeepSeek
deepseek-v4-flash
DeepSeek
Google
gemini-3-7-flash
Meta
muse-spark-1-2
Meta
Moonshot
kimi-k3
Moonshot
kimi-k2-7-code
Moonshot
OpenAI
gpt-5-6-sol
OpenAI
gpt-5-6-terra
OpenAI
gpt-5-6-luna
OpenAI
xAI
grok-4-6
xAI
Zhipu
glm-5-3
Zhipu