Model Bench

Which model and effort level gives a coding agent the most capability per dollar, scored for the harness you actually run.

Data as of
Models
6, 24 configurations
Benchmarks
5, 75 measurements
Method
How it works, sources
Colour theme

Most capability per dollar with the harnesses you can run at work, at α 6

in OpenCode

Capability
79.7
Cost per task
$0.39
Value
87.8
Confidence
0.63

Close behind

A doubling of cost must buy 6 capability points to break even.

Harness

Claude Code for Anthropic models, OpenCode for the rest. What you can run at work.

Least certain setting. None of these models has a direct OpenCode run; the guess rests on one 38-task comparison and one out-of-list model on Terminal-Bench.

Grok discount

Temporary 90% off. Grok pays 10% of list cost.

Show

Front within 15 value points of the best, the best per cost band, and free models.

Capability against cost

  • Anthropic
  • OpenAI
  • xAI
  • Alibaba
  • Pareto front
  • Equal value to the best
  • Effort levels of one model

Faint points are hidden by the current view. Anything above the red dashed line would beat GPT-6.1 Sol xhigh; the line rises 6 points for each doubling of cost. Hover or click a point, or focus the chart and use the arrow keys.

Selected GPT-6.1 Sol xhigh, capability 79.7, $0.39 per task, value 87.8.

Ranking

Click a row to trace its numbers. Sort by any column.

Configurations ranked by value. Upright numbers were measured for this configuration; italic numbers were adjusted, with a footnote letter for each adjustment.
ResultBenchmark scores used, raw
1bestfrontxhighOpenCode79.787.8$0.3951.073.7eo51.0o63.8eop54.7e530.63
2fronthighOpenCode76.786.6$0.3250.072.2o48.5o62.3eop53.2e500.65
3frontmaxOpenCode82.685.5$0.7252.075.2eo53.0o65.3eop56.2e560.63
4frontmediumOpenCode70.784.2$0.2148.069.2eo45.0o59.3op50.2490.66
5highClaude Code81.075.9$1.8254.062.4e57.1e59.3hp59.1e770.76
6frontxhighClaude Code85.875.1$3.4656.065.4e60.1e62.3ehp62.1e810.75
7frontmaxClaude Code89.473.9$5.9858.068.463.165.3ehp65.1e960.83
17frontxhighOpenCode34.262.0$0.0435.059.1eo5.0o48.1eop37.9e1160.63
22freedefaultOpenCode46.955.5$0.3740.055.7o22.3o––560.51

Showing 9 of 24 configurations (Sweet spot).

Upright numbers were measured for this exact model, effort and harness. Italic numbers were adjusted by the engine; the red letters say how. The record sheet lists every adjustment with its size.

e
scaled from another effort level
h
converted from another harness
o
OpenCode offset applied
p
previous-generation proxy
est
estimated by the data author
d
Grok discount applied to cost
–
no run; the other benchmarks share its weight
OpenCode
in italics: no direct runs, estimated with the offset

How the numbers are made

Public leaderboards disagree because they test different things, run models in different agent harnesses, and rarely show cost. Model Bench puts five benchmarks on one 0 to 100 capability scale, recomputes that scale for the harness you would actually run, and ranks every model and effort level by capability per dollar. Data as of 9 Oct 2026. Every value in the ranking traces back to a source listed at the bottom of this page.

  1. Five benchmarks, weighted by what they tell you

    Each benchmark covers something the others miss. Weights add up to 1. The AA index weighs most because it is the only source measured at every effort level, so it anchors how a model changes with effort. DeepSWE and Terminal-Bench are long agent tasks, close to real coding-agent work. SWE-rebench uses fresh issues but every value here comes from the previous model generation, so it counts less. FrontierCode judges code quality but covers the fewest models.

    BenchmarkWeightRaw score at 0Raw score at 100Harness mattersWhat it measures
    AA Intelligence Index v4.3.20.281560noTen evaluations run independently by Artificial Analysis. The only source that covers every effort level.
    DeepSWE v1.10.223080yes113 long-horizon tasks written from scratch across 91 repositories. Contamination-free, common mini-swe-agent harness.
    Terminal-Bench 4.00.20570yes66 tasks of agentic work in a terminal. This is where models differ most.
    SWE-rebench rolling0.182070yesNebius. Fresh GitHub issues in a rolling window, 100+ models, fixed scaffold, five runs per task. Current window ends 2026-07-01, so every model here carries its predecessor's value.
    FrontierCode v1.10.123060noCognition. Quality of production code. Each model runs in its vendor's harness, and the board covers the fewest models, so it gets the lowest weight.
  2. Normalize each score to 0 to 100

    Raw scores live on different scales, so each benchmark gets two anchors: the raw score that maps to 0 and the one that maps to 100 (table above). Scores outside the anchors are clipped.

    Example: 56 on Terminal-Bench becomes 78.5, because its anchors are 5 and 70.

    normalized = (raw − low anchor) ÷ (high anchor − low anchor) × 100

  3. Fill the gaps, and trust filled values less

    Few configurations were measured on every benchmark, at every effort level, in every harness. For each benchmark the engine takes the measurement closest to what you asked for, then adjusts it. Each adjustment multiplies that value's reliability, so it carries less weight in the final score. The table marks every adjusted value in italics with a letter, and the record sheet lists the exact adjustments.

    Harness first
    It looks for a run in the harness you picked: the vendor's own harness in Native mode, the minimal common harness in Common mode. None of these models has an OpenCode run, so in Company mode non-Anthropic models start from common-harness runs.
    Effort scaling e
    Among the candidate runs it picks the effort level whose AA index is closest. If that is not the effort you are looking at, the score moves 1.5 benchmark points for each point of AA index between the two. Reliability × 0.8.
    Previous-generation proxies p
    Where a model has not been run yet, its predecessor's score stands in, with a reliability of 0.4 to 0.5 set per measurement. Every proxy is labelled in the sources table.
    Harness conversion h
    If the only run is in the other harness, the vendor's typical gap is added or subtracted (the conversion table). Only Terminal-Bench, DeepSWE and SWE-rebench are converted; the AA index and FrontierCode do not depend on the harness. Reliability × 0.6.
    OpenCode offset o
    In Company mode, OpenAI, xAI and Alibaba models run in OpenCode. Their harness-sensitive scores get the offset you set (default −3). This is the least certain number in the app: the only direct OpenCode measurement is of a model outside this list (GLM-5.3, about 2 points below the common harness on Terminal-Bench). Reliability × 0.7.
    Vendor harness gain over the common harness, in raw benchmark points
    VendorVendor harnessTerminal-BenchDeepSWESWE-rebench
    AnthropicClaude Code+2.4−3.0−4.1
    OpenAICodex0.0−3.5−4.3
    xAIGrok Build+11.6+6.0+6.0
    AlibabaQwen Code0.00.00.0
  4. Combine into capability and confidence

    A benchmark with no measurement drops out and the others share its weight. Confidence is the denominator, Σ (weight × reliability). It is 1.00 only when all five benchmarks were measured directly for that model, effort and harness. Lower means more of the score rests on filled-in values. Qwen3.8 Flash-Next has only the AA index, so its confidence is 0.51: its capability is the AA index alone.

    capability = Σ (weight × reliability × normalized) ÷ Σ (weight × reliability)

    confidence = Σ (weight × reliability)

  5. Rank by value

    α says how many capability points a doubling of cost is worth. At 0, cost is ignored. At the default 6, a configuration that costs twice as much has to score 6 points higher to break even. At 16, cost dominates and cheap models win. Because cost is on a log axis in the chart, every line of equal value is straight; the dashed red line is the one through the current best, so nothing sits above it.

    A configuration is on the Pareto front when no other one is at least as cheap and at least as capable, and strictly better on one of the two. The Sweet spot view shows front configurations within 15 value points of the best, the best-value configuration in each cost band (under $0.05, $0.05 to $0.25, $0.25 to $1.00, $1.00 and up), and anything free to you. Faint points in the chart are the ones the view hides.

    Cost per task, output speed and time to first token are stored per configuration; values marked est. are estimates. With the Grok discount on, Grok's cost is multiplied by 0.1. Cost does not depend on the harness here, though in practice it does (see the open gaps).

    value = capability − α × log2(cost per task in USD)

  6. Harness evidence

    Cases where the same model was run on the same benchmark, at the same effort, in both a minimal common harness and its vendor's own harness. The vendor harness does not always help: Grok gains a lot in Grok Build, Anthropic and OpenAI models move a few points either way on Terminal-Bench, and both SWE-rebench pairs score lower in the vendor harness. The Terminal-Bench and SWE-rebench columns of the conversion table in step 3 are the averages of these pairs; the DeepSWE column is an estimate. Every model here except Qwen now has runs in both harnesses on Terminal-Bench and DeepSWE, so the table mostly matters for SWE-rebench.

    Model and benchmark020406080Vendor harness gain
    Opus 5.5 max · Terminal-Bench 4.0+3.5 in Claude Code
    Sonnet 5.5 max · Terminal-Bench 4.0+2.6 in Claude Code
    Sonnet 5.5 xhigh · Terminal-Bench 4.0+1.0 in Claude Code
    GPT-6.1 Sol xhigh · Terminal-Bench 4.0+0.5 in Codex
    GPT-6.1 Sol max · Terminal-Bench 4.0−3.0 in Codex
    GPT-6 Luna max · Terminal-Bench 4.0+2.6 in Codex
    Grok 4.7 xhigh · Terminal-Bench 4.0+11.6 in Grok Build
    Fable 5 high · SWE-rebench−4.1 in Claude Code
    GPT-5.6 Sol medium · SWE-rebench−4.3 in Codex
    common minimal harness vendor harness
  7. Open data gaps

    Known weaknesses in the current data. Read the ranking with these in mind.

    • Datacurve's own DeepSWE board lists none of these models. DeepSWE values come from Artificial Analysis runs in each vendor's harness (Coding Agent Index) and from vendor self-reports; Opus 5.5 below max in the common harness still carries Opus 5 values.
    • SWE-rebench's window still ends 2026-07-01. All six models carry previous-generation values at reliability 0.4.
    • OpenCode is measured directly for only one model, which is not in this list (GLM-5.3 on Artificial Analysis): about 2 points below the common harness on Terminal-Bench and about 8 below on DeepSWE, against a different runner. The −3 default rests on that and on one 38-task comparison.
    • Cost per task is the Artificial Analysis Intelligence Index cost at list price. Agent harnesses cost more per task: on AA's Coding Agent Index the ratio is about 1.3× for Sonnet 5.5 and 3.3× for GPT-6.1 Sol at medium, and one test showed Claude Code at about 6× the cost of other CLIs on the same model.
    • FrontierCode runs each model in its vendor's harness (Claude Code, Codex, Grok Build), so it is not fully harness-neutral. The app treats it as neutral.
    • Speed and time to first token are Artificial Analysis rolling medians on the data date and drift from week to week. GPT-6 Luna medium has no current speed measurement; its value is from launch.

Sources

All 75 benchmark measurements behind the ranking, by model. AA index values for every effort level come from Artificial Analysis. Benchmark home pages: AA Intelligence Index, DeepSWE, Terminal-Bench, SWE-rebench, FrontierCode.

BenchmarkHarnessEffortRaw scoreReliabilitySource
GPT-6.1 Sol
Terminal-Benchcommon harnesslow30.801.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessmedium48.01.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnesshigh51.501.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessxhigh54.01.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessmax56.01.0AA, mini-swe-agent artificialanalysis.ai
Terminal-BenchCodexlow49.01.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
Terminal-BenchCodexmedium51.501.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
Terminal-BenchCodexhigh50.01.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
Terminal-BenchCodexxhigh54.501.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
Terminal-BenchCodexmax53.01.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
Terminal-BenchCodexmax58.181.0Terminal-Bench official board, Codex tbench.ai
DeepSWEcommon harnesshigh75.220.6OpenAI self-reported llm-stats.com
DeepSWECodexlow67.601.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
DeepSWECodexmedium72.01.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
DeepSWECodexhigh70.501.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
DeepSWECodexxhigh73.201.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
DeepSWECodexmax69.601.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
SWE-rebenchcommon harnessmedium62.300.4proxy: GPT-5.6 Sol [medium] swe-rebench.com
FrontierCodeharness-neutralmedium50.201.0FrontierCode v1.1 Main (Cognition, run in Codex) cognition.com
Claude Opus 5.5
Terminal-Benchcommon harnesslow31.301.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessmedium52.501.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnesshigh57.01.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessxhigh59.601.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessmax59.601.0AA, mini-swe-agent artificialanalysis.ai
Terminal-BenchClaude Codemax63.101.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
Terminal-BenchClaude Codemax64.851.0Terminal-Bench official board, Claude Code tbench.ai
DeepSWEcommon harnesslow58.100.5proxy: Opus 5 (Epoch data via llmrun) llmrun.dev
DeepSWEcommon harnessmedium68.900.5proxy: Opus 5 (Epoch data via llmrun) llmrun.dev
DeepSWEcommon harnesshigh72.800.5proxy: Opus 5 (Epoch data via llmrun) llmrun.dev
DeepSWEcommon harnessxhigh73.200.5proxy: Opus 5 (Epoch data via llmrun) llmrun.dev
DeepSWEcommon harnessmax74.200.6Anthropic self-reported (system card, mean of 5 trials, harness not stated) anthropic.com
DeepSWEClaude Codemax68.401.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
SWE-rebenchcommon harnesshigh63.400.4proxy: Opus 5 [high] swe-rebench.com
FrontierCodeharness-neutralmedium54.601.0FrontierCode v1.1 Main (Cognition, run in Claude Code) cognition.com
Claude Sonnet 5.5
Terminal-Benchcommon harnesslow20.701.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessmedium29.801.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnesshigh43.901.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessxhigh57.101.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessmax63.601.0AA, mini-swe-agent artificialanalysis.ai
Terminal-BenchClaude Codelow25.301.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
Terminal-BenchClaude Codemedium27.301.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
Terminal-BenchClaude Codehigh41.901.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
Terminal-BenchClaude Codexhigh58.101.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
Terminal-BenchClaude Codemax66.201.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
Terminal-BenchClaude Codemax61.821.0Terminal-Bench official board, Claude Code tbench.ai
DeepSWEcommon harnessmax71.00.6Anthropic self-reported (system card, mean of 5 trials, harness not stated) anthropic.com
DeepSWEClaude Codelow61.901.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
DeepSWEClaude Codemedium65.501.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
DeepSWEClaude Codehigh66.701.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
DeepSWEClaude Codexhigh68.401.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
DeepSWEClaude Codemax72.01.0AA Coding Agent Index v1.5, Claude Code artificialanalysis.ai
SWE-rebenchcommon harnesshigh56.800.4proxy: Sonnet 5 [high] swe-rebench.com
FrontierCodeharness-neutralxhigh52.101.0FrontierCode v1.1 Main (Cognition, run in Claude Code) cognition.com
FrontierCodeharness-neutralmax46.200.8FrontierCode v1.1 Main, Cognition run reported in Anthropic's system card anthropic.com
GPT-6 Luna
Terminal-Benchcommon harnesslow0.01.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessmedium2.501.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnesshigh4.501.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessxhigh8.01.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessmax12.601.0AA, mini-swe-agent artificialanalysis.ai
Terminal-BenchCodexmax15.201.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
Terminal-BenchCodexmax16.361.0Terminal-Bench official board, Codex tbench.ai
DeepSWEcommon harnessmax66.600.6OpenAI self-reported (not on Datacurve's board) codingfleet.com
DeepSWECodexmax63.701.0AA Coding Agent Index v1.5, Codex artificialanalysis.ai
SWE-rebenchcommon harnessmedium43.600.4proxy: GPT-5.6 Luna [medium] swe-rebench.com
FrontierCodeharness-neutralmax42.401.0FrontierCode v1.1 Main (Cognition, run in Codex) cognition.com
Grok 4.7
Terminal-Benchcommon harnesslow16.201.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnesshigh25.01.0AA, mini-swe-agent artificialanalysis.ai
Terminal-Benchcommon harnessxhigh26.01.0AA, mini-swe-agent artificialanalysis.ai
Terminal-BenchGrok Buildxhigh37.581.0Terminal-Bench official board, Grok Build tbench.ai
DeepSWEGrok Buildxhigh72.601.0AA Coding Agent Index v1.5, Grok Build artificialanalysis.ai
DeepSWEcommon harnessxhigh66.700.5proxy: Grok 4.6 [xhigh] (Epoch data via llmrun) llmrun.dev
SWE-rebenchcommon harnesshigh63.800.4proxy: Grok 4.5 [high] swe-rebench.com
FrontierCodeharness-neutralhigh47.601.0FrontierCode v1.1 Main (Cognition, run in Grok Build) cognition.com
Qwen3.8 Flash-Next
Terminal-Benchcommon harnessdefault25.301.0AA, mini-swe-agent artificialanalysis.ai
DeepSWEcommon harnessdefault58.700.6Qwen self-reported (model card) huggingface.co