Most capability per dollar with the harnesses you can run at work, at α 6
in OpenCode
- Capability
- 79.7
- Cost per task
- $0.39
- Value
- 87.8
- Confidence
- 0.63
Close behind
Capability against cost
- Anthropic
- OpenAI
- xAI
- Alibaba
- Pareto front
- Equal value to the best
- Effort levels of one model
Faint points are hidden by the current view. Anything above the red dashed line would beat GPT-6.1 Sol xhigh; the line rises 6 points for each doubling of cost. Hover or click a point, or focus the chart and use the arrow keys.
Selected GPT-6.1 Sol xhigh, capability 79.7, $0.39 per task, value 87.8.
Ranking
Click a row to trace its numbers. Sort by any column.
| Result | Benchmark scores used, raw | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | bestfront | xhigh | OpenCode | 79.7 | 87.8 | $0.39 | 51.0 | 73.7eo | 51.0o | 63.8eop | 54.7e | 53 | 0.63 |
| 2 | front | high | OpenCode | 76.7 | 86.6 | $0.32 | 50.0 | 72.2o | 48.5o | 62.3eop | 53.2e | 50 | 0.65 |
| 3 | front | max | OpenCode | 82.6 | 85.5 | $0.72 | 52.0 | 75.2eo | 53.0o | 65.3eop | 56.2e | 56 | 0.63 |
| 4 | front | medium | OpenCode | 70.7 | 84.2 | $0.21 | 48.0 | 69.2eo | 45.0o | 59.3op | 50.2 | 49 | 0.66 |
| 5 | high | Claude Code | 81.0 | 75.9 | $1.82 | 54.0 | 62.4e | 57.1e | 59.3hp | 59.1e | 77 | 0.76 | |
| 6 | front | xhigh | Claude Code | 85.8 | 75.1 | $3.46 | 56.0 | 65.4e | 60.1e | 62.3ehp | 62.1e | 81 | 0.75 |
| 7 | front | max | Claude Code | 89.4 | 73.9 | $5.98 | 58.0 | 68.4 | 63.1 | 65.3ehp | 65.1e | 96 | 0.83 |
| 17 | front | xhigh | OpenCode | 34.2 | 62.0 | $0.04 | 35.0 | 59.1eo | 5.0o | 48.1eop | 37.9e | 116 | 0.63 |
| 22 | free | default | OpenCode | 46.9 | 55.5 | $0.37 | 40.0 | 55.7o | 22.3o | – | – | 56 | 0.51 |
Showing 9 of 24 configurations (Sweet spot).
Upright numbers were measured for this exact model, effort and harness. Italic numbers were adjusted by the engine; the red letters say how. The record sheet lists every adjustment with its size.
- e
- scaled from another effort level
- h
- converted from another harness
- o
- OpenCode offset applied
- p
- previous-generation proxy
- est
- estimated by the data author
- d
- Grok discount applied to cost
- –
- no run; the other benchmarks share its weight
- OpenCode
- in italics: no direct runs, estimated with the offset
How the numbers are made
Public leaderboards disagree because they test different things, run models in different agent harnesses, and rarely show cost. Model Bench puts five benchmarks on one 0 to 100 capability scale, recomputes that scale for the harness you would actually run, and ranks every model and effort level by capability per dollar. Data as of 9 Oct 2026. Every value in the ranking traces back to a source listed at the bottom of this page.
Five benchmarks, weighted by what they tell you
Each benchmark covers something the others miss. Weights add up to 1. The AA index weighs most because it is the only source measured at every effort level, so it anchors how a model changes with effort. DeepSWE and Terminal-Bench are long agent tasks, close to real coding-agent work. SWE-rebench uses fresh issues but every value here comes from the previous model generation, so it counts less. FrontierCode judges code quality but covers the fewest models.
Benchmark Weight Raw score at 0 Raw score at 100 Harness matters What it measures AA Intelligence Index v4.3.2 0.28 15 60 no Ten evaluations run independently by Artificial Analysis. The only source that covers every effort level. DeepSWE v1.1 0.22 30 80 yes 113 long-horizon tasks written from scratch across 91 repositories. Contamination-free, common mini-swe-agent harness. Terminal-Bench 4.0 0.20 5 70 yes 66 tasks of agentic work in a terminal. This is where models differ most. SWE-rebench rolling 0.18 20 70 yes Nebius. Fresh GitHub issues in a rolling window, 100+ models, fixed scaffold, five runs per task. Current window ends 2026-07-01, so every model here carries its predecessor's value. FrontierCode v1.1 0.12 30 60 no Cognition. Quality of production code. Each model runs in its vendor's harness, and the board covers the fewest models, so it gets the lowest weight. Normalize each score to 0 to 100
Raw scores live on different scales, so each benchmark gets two anchors: the raw score that maps to 0 and the one that maps to 100 (table above). Scores outside the anchors are clipped.
Example: 56 on Terminal-Bench becomes 78.5, because its anchors are 5 and 70.
normalized = (raw − low anchor) ÷ (high anchor − low anchor) × 100
Fill the gaps, and trust filled values less
Few configurations were measured on every benchmark, at every effort level, in every harness. For each benchmark the engine takes the measurement closest to what you asked for, then adjusts it. Each adjustment multiplies that value's reliability, so it carries less weight in the final score. The table marks every adjusted value in italics with a letter, and the record sheet lists the exact adjustments.
- Harness first
- It looks for a run in the harness you picked: the vendor's own harness in Native mode, the minimal common harness in Common mode. None of these models has an OpenCode run, so in Company mode non-Anthropic models start from common-harness runs.
- Effort scaling e
- Among the candidate runs it picks the effort level whose AA index is closest. If that is not the effort you are looking at, the score moves 1.5 benchmark points for each point of AA index between the two. Reliability × 0.8.
- Previous-generation proxies p
- Where a model has not been run yet, its predecessor's score stands in, with a reliability of 0.4 to 0.5 set per measurement. Every proxy is labelled in the sources table.
- Harness conversion h
- If the only run is in the other harness, the vendor's typical gap is added or subtracted (the conversion table). Only Terminal-Bench, DeepSWE and SWE-rebench are converted; the AA index and FrontierCode do not depend on the harness. Reliability × 0.6.
- OpenCode offset o
- In Company mode, OpenAI, xAI and Alibaba models run in OpenCode. Their harness-sensitive scores get the offset you set (default −3). This is the least certain number in the app: the only direct OpenCode measurement is of a model outside this list (GLM-5.3, about 2 points below the common harness on Terminal-Bench). Reliability × 0.7.
Vendor harness gain over the common harness, in raw benchmark points Vendor Vendor harness Terminal-Bench DeepSWE SWE-rebench Anthropic Claude Code +2.4 −3.0 −4.1 OpenAI Codex 0.0 −3.5 −4.3 xAI Grok Build +11.6 +6.0 +6.0 Alibaba Qwen Code 0.0 0.0 0.0 Combine into capability and confidence
A benchmark with no measurement drops out and the others share its weight. Confidence is the denominator, Σ (weight × reliability). It is 1.00 only when all five benchmarks were measured directly for that model, effort and harness. Lower means more of the score rests on filled-in values. Qwen3.8 Flash-Next has only the AA index, so its confidence is 0.51: its capability is the AA index alone.
capability = Σ (weight × reliability × normalized) ÷ Σ (weight × reliability)
confidence = Σ (weight × reliability)
Rank by value
α says how many capability points a doubling of cost is worth. At 0, cost is ignored. At the default 6, a configuration that costs twice as much has to score 6 points higher to break even. At 16, cost dominates and cheap models win. Because cost is on a log axis in the chart, every line of equal value is straight; the dashed red line is the one through the current best, so nothing sits above it.
A configuration is on the Pareto front when no other one is at least as cheap and at least as capable, and strictly better on one of the two. The Sweet spot view shows front configurations within 15 value points of the best, the best-value configuration in each cost band (under $0.05, $0.05 to $0.25, $0.25 to $1.00, $1.00 and up), and anything free to you. Faint points in the chart are the ones the view hides.
Cost per task, output speed and time to first token are stored per configuration; values marked est. are estimates. With the Grok discount on, Grok's cost is multiplied by 0.1. Cost does not depend on the harness here, though in practice it does (see the open gaps).
value = capability − α × log2(cost per task in USD)
Harness evidence
Cases where the same model was run on the same benchmark, at the same effort, in both a minimal common harness and its vendor's own harness. The vendor harness does not always help: Grok gains a lot in Grok Build, Anthropic and OpenAI models move a few points either way on Terminal-Bench, and both SWE-rebench pairs score lower in the vendor harness. The Terminal-Bench and SWE-rebench columns of the conversion table in step 3 are the averages of these pairs; the DeepSWE column is an estimate. Every model here except Qwen now has runs in both harnesses on Terminal-Bench and DeepSWE, so the table mostly matters for SWE-rebench.
Model and benchmark020406080Vendor harness gainOpus 5.5 max · Terminal-Bench 4.0+3.5 in Claude CodeSonnet 5.5 max · Terminal-Bench 4.0+2.6 in Claude CodeSonnet 5.5 xhigh · Terminal-Bench 4.0+1.0 in Claude CodeGPT-6.1 Sol xhigh · Terminal-Bench 4.0+0.5 in CodexGPT-6.1 Sol max · Terminal-Bench 4.0−3.0 in CodexGPT-6 Luna max · Terminal-Bench 4.0+2.6 in CodexGrok 4.7 xhigh · Terminal-Bench 4.0+11.6 in Grok BuildFable 5 high · SWE-rebench−4.1 in Claude CodeGPT-5.6 Sol medium · SWE-rebench−4.3 in Codexcommon minimal harness vendor harness Open data gaps
Known weaknesses in the current data. Read the ranking with these in mind.
- Datacurve's own DeepSWE board lists none of these models. DeepSWE values come from Artificial Analysis runs in each vendor's harness (Coding Agent Index) and from vendor self-reports; Opus 5.5 below max in the common harness still carries Opus 5 values.
- SWE-rebench's window still ends 2026-07-01. All six models carry previous-generation values at reliability 0.4.
- OpenCode is measured directly for only one model, which is not in this list (GLM-5.3 on Artificial Analysis): about 2 points below the common harness on Terminal-Bench and about 8 below on DeepSWE, against a different runner. The −3 default rests on that and on one 38-task comparison.
- Cost per task is the Artificial Analysis Intelligence Index cost at list price. Agent harnesses cost more per task: on AA's Coding Agent Index the ratio is about 1.3× for Sonnet 5.5 and 3.3× for GPT-6.1 Sol at medium, and one test showed Claude Code at about 6× the cost of other CLIs on the same model.
- FrontierCode runs each model in its vendor's harness (Claude Code, Codex, Grok Build), so it is not fully harness-neutral. The app treats it as neutral.
- Speed and time to first token are Artificial Analysis rolling medians on the data date and drift from week to week. GPT-6 Luna medium has no current speed measurement; its value is from launch.
Sources
All 75 benchmark measurements behind the ranking, by model. AA index values for every effort level come from Artificial Analysis. Benchmark home pages: AA Intelligence Index, DeepSWE, Terminal-Bench, SWE-rebench, FrontierCode.
| Benchmark | Harness | Effort | Raw score | Reliability | Source |
|---|---|---|---|---|---|
| GPT-6.1 Sol | |||||
| Terminal-Bench | common harness | low | 30.80 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | medium | 48.0 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | high | 51.50 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | xhigh | 54.0 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | max | 56.0 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | Codex | low | 49.0 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| Terminal-Bench | Codex | medium | 51.50 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| Terminal-Bench | Codex | high | 50.0 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| Terminal-Bench | Codex | xhigh | 54.50 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| Terminal-Bench | Codex | max | 53.0 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| Terminal-Bench | Codex | max | 58.18 | 1.0 | Terminal-Bench official board, Codex tbench.ai |
| DeepSWE | common harness | high | 75.22 | 0.6 | OpenAI self-reported llm-stats.com |
| DeepSWE | Codex | low | 67.60 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| DeepSWE | Codex | medium | 72.0 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| DeepSWE | Codex | high | 70.50 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| DeepSWE | Codex | xhigh | 73.20 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| DeepSWE | Codex | max | 69.60 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| SWE-rebench | common harness | medium | 62.30 | 0.4 | proxy: GPT-5.6 Sol [medium] swe-rebench.com |
| FrontierCode | harness-neutral | medium | 50.20 | 1.0 | FrontierCode v1.1 Main (Cognition, run in Codex) cognition.com |
| Claude Opus 5.5 | |||||
| Terminal-Bench | common harness | low | 31.30 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | medium | 52.50 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | high | 57.0 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | xhigh | 59.60 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | max | 59.60 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | Claude Code | max | 63.10 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| Terminal-Bench | Claude Code | max | 64.85 | 1.0 | Terminal-Bench official board, Claude Code tbench.ai |
| DeepSWE | common harness | low | 58.10 | 0.5 | proxy: Opus 5 (Epoch data via llmrun) llmrun.dev |
| DeepSWE | common harness | medium | 68.90 | 0.5 | proxy: Opus 5 (Epoch data via llmrun) llmrun.dev |
| DeepSWE | common harness | high | 72.80 | 0.5 | proxy: Opus 5 (Epoch data via llmrun) llmrun.dev |
| DeepSWE | common harness | xhigh | 73.20 | 0.5 | proxy: Opus 5 (Epoch data via llmrun) llmrun.dev |
| DeepSWE | common harness | max | 74.20 | 0.6 | Anthropic self-reported (system card, mean of 5 trials, harness not stated) anthropic.com |
| DeepSWE | Claude Code | max | 68.40 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| SWE-rebench | common harness | high | 63.40 | 0.4 | proxy: Opus 5 [high] swe-rebench.com |
| FrontierCode | harness-neutral | medium | 54.60 | 1.0 | FrontierCode v1.1 Main (Cognition, run in Claude Code) cognition.com |
| Claude Sonnet 5.5 | |||||
| Terminal-Bench | common harness | low | 20.70 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | medium | 29.80 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | high | 43.90 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | xhigh | 57.10 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | max | 63.60 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | Claude Code | low | 25.30 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| Terminal-Bench | Claude Code | medium | 27.30 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| Terminal-Bench | Claude Code | high | 41.90 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| Terminal-Bench | Claude Code | xhigh | 58.10 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| Terminal-Bench | Claude Code | max | 66.20 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| Terminal-Bench | Claude Code | max | 61.82 | 1.0 | Terminal-Bench official board, Claude Code tbench.ai |
| DeepSWE | common harness | max | 71.0 | 0.6 | Anthropic self-reported (system card, mean of 5 trials, harness not stated) anthropic.com |
| DeepSWE | Claude Code | low | 61.90 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| DeepSWE | Claude Code | medium | 65.50 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| DeepSWE | Claude Code | high | 66.70 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| DeepSWE | Claude Code | xhigh | 68.40 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| DeepSWE | Claude Code | max | 72.0 | 1.0 | AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai |
| SWE-rebench | common harness | high | 56.80 | 0.4 | proxy: Sonnet 5 [high] swe-rebench.com |
| FrontierCode | harness-neutral | xhigh | 52.10 | 1.0 | FrontierCode v1.1 Main (Cognition, run in Claude Code) cognition.com |
| FrontierCode | harness-neutral | max | 46.20 | 0.8 | FrontierCode v1.1 Main, Cognition run reported in Anthropic's system card anthropic.com |
| GPT-6 Luna | |||||
| Terminal-Bench | common harness | low | 0.0 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | medium | 2.50 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | high | 4.50 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | xhigh | 8.0 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | max | 12.60 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | Codex | max | 15.20 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| Terminal-Bench | Codex | max | 16.36 | 1.0 | Terminal-Bench official board, Codex tbench.ai |
| DeepSWE | common harness | max | 66.60 | 0.6 | OpenAI self-reported (not on Datacurve's board) codingfleet.com |
| DeepSWE | Codex | max | 63.70 | 1.0 | AA Coding Agent Index v1.5, Codex artificialanalysis.ai |
| SWE-rebench | common harness | medium | 43.60 | 0.4 | proxy: GPT-5.6 Luna [medium] swe-rebench.com |
| FrontierCode | harness-neutral | max | 42.40 | 1.0 | FrontierCode v1.1 Main (Cognition, run in Codex) cognition.com |
| Grok 4.7 | |||||
| Terminal-Bench | common harness | low | 16.20 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | high | 25.0 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | common harness | xhigh | 26.0 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| Terminal-Bench | Grok Build | xhigh | 37.58 | 1.0 | Terminal-Bench official board, Grok Build tbench.ai |
| DeepSWE | Grok Build | xhigh | 72.60 | 1.0 | AA Coding Agent Index v1.5, Grok Build artificialanalysis.ai |
| DeepSWE | common harness | xhigh | 66.70 | 0.5 | proxy: Grok 4.6 [xhigh] (Epoch data via llmrun) llmrun.dev |
| SWE-rebench | common harness | high | 63.80 | 0.4 | proxy: Grok 4.5 [high] swe-rebench.com |
| FrontierCode | harness-neutral | high | 47.60 | 1.0 | FrontierCode v1.1 Main (Cognition, run in Grok Build) cognition.com |
| Qwen3.8 Flash-Next | |||||
| Terminal-Bench | common harness | default | 25.30 | 1.0 | AA, mini-swe-agent artificialanalysis.ai |
| DeepSWE | common harness | default | 58.70 | 0.6 | Qwen self-reported (model card) huggingface.co |