← back to the article

Which model to run

Run Grok 4.5 for everyday work: Blender 5.2 API, IK solver behavior, coordinate spaces. Reach for Claude Sonnet 5 on skinning and deformation. No model clears the bar yet on image-to-3D pipeline; review the checks. measurement design has no floor tests yet, so the bar can't be evaluated there - a suite coverage gap, not a model failure.

Headline follows your cheapest lens. The two frontier charts and the per-service table below show both cost and latency.

Cheapest and fastest picks differ on: Blender 5.2 API (cheapest Grok 4.5, fastest Claude Sonnet 5); IK solver behavior (cheapest Grok 4.5, fastest Claude Sonnet 5); coordinate spaces (cheapest Grok 4.5, fastest Claude Sonnet 5).

Benchmark comparison

Benchmark: Claude Sonnet 5. Positive cost/latency is worse (more expensive/slower than the benchmark); positive disc is better (higher quality).

modelcost vs benchmarklatency vs benchmarkdisc vs benchmark
Grok 4.5-30%+58%-0.09
GPT 5.6 Sol+56%+28%-0.02
Gemini 3.1 Pro Preview+140%+91%+0.05

Cost and latency vs quality

cost vs quality
64%73%82%91%100%$0.00$0.75$1.49$2.24$2.98cost per 100 tests (USD), lower is betterfloor pass-rate
latency vs quality
64%73%82%91%100%0.000s7.183s14.367s21.550s28.733smedian latency per test (s), lower is betterfloor pass-rate
Claude Sonnet 5 (you are here)GPT 5.6 SolGemini 3.1 Pro PreviewGrok 4.5

All models at a glance

floormust-pass correctness and guardrails, the routing gate (you want 100%).
discdiscriminating, the graded hard-reasoning score, 0 to 1, used to rank models.
latencymedian response time.
cost /100API cost per 100 tests measured on this suite - it reflects each model's own verbosity here, not just its list price, so it moves if your real tasks are longer (per-test is fractions of a cent; the run-cost panel shows real totals).

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Sonnet 571%0.7713.1s$1.08you are here
GPT 5.6 Sol71%0.7516.8s$1.68
Gemini 3.1 Pro Preview64%0.8225.0s$2.59
Grok 4.564%0.6820.6s$0.76

Per-service routing

servicecheapestfastest
Blender 5.2 APIGrok 4.5 $0.95/100Claude Sonnet 5 26.9s
IK solver behaviorGrok 4.5 $0.96/100Claude Sonnet 5 13.4s
image-to-3D pipelineno model clears the bar yet (GPT 5.6 Sol is closest)
measurement designno floor tests in this category - by discriminating score, Gemini 3.1 Pro Preview ranks highest
skinning and deformationClaude Sonnet 5 cheapest and fastest $1.38/100, 19.7s
coordinate spacesGrok 4.5 $0.63/100Claude Sonnet 5 9.2s

Per-category detail

Blender 5.2 API -> run Grok 4.5(cheapest that clears the bar)

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Sonnet 5100%n/a26.9s$2.38you are here
Gemini 3.1 Pro Preview100%n/a43.5s$4.63
GPT 5.6 Sol100%n/a34.9s$2.42
Grok 4.5100%n/a27.3s$0.95run this

Every model cleared every floor test in this category.

IK solver behavior -> run Grok 4.5(cheapest that clears the bar)

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Sonnet 5100%0.8313.4s$1.02you are here
Gemini 3.1 Pro Preview100%0.8628.9s$2.38
GPT 5.6 Sol100%0.8429.5s$2.43
Grok 4.5100%0.7337.7s$0.96run this

Every model cleared every floor test in this category.

image-to-3D pipeline -> NO MODEL CLEARS THE BAR - no model clears the bar yet, openai:gpt-5.6-sol is closest, review the checks

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Gemini 3.1 Pro Preview50%n/a21.0s$2.24
GPT 5.6 Sol50%n/a14.0s$1.19
Claude Sonnet 533%n/a12.4s$1.02you are here
Grok 4.533%n/a22.4s$0.66
Floor tests failed here (deduped across repeats)
  • Claude Sonnet 5 fails An asymmetric detail flips sides between the two views (1/1 runs)
  • Claude Sonnet 5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Claude Sonnet 5 fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Claude Sonnet 5 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Gemini 3.1 Pro Preview fails An asymmetric detail flips sides between the two views (1/1 runs)
  • Gemini 3.1 Pro Preview fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Gemini 3.1 Pro Preview fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • GPT 5.6 Sol fails An asymmetric detail flips sides between the two views (1/1 runs)
  • GPT 5.6 Sol fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • GPT 5.6 Sol fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Grok 4.5 fails An asymmetric detail flips sides between the two views (1/1 runs)
  • Grok 4.5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Grok 4.5 fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Grok 4.5 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
measurement design -> NO FLOOR TESTS HERE - this category has no floor tests, so the bar can't be evaluated here; by discriminating score, google:gemini-3.1-pro-preview ranks highest

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Sonnet 5n/a0.8311.2s$0.91you are here
Gemini 3.1 Pro Previewn/a0.8428.1s$2.73
GPT 5.6 Soln/a0.7816.0s$1.32
Grok 4.5n/a0.7019.6s$0.80

Every model cleared every floor test in this category.

skinning and deformation -> run Claude Sonnet 5(the only model that clears the bar)

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Sonnet 5100%0.7319.7s$1.38run this
Gemini 3.1 Pro Preview0%0.8226.0s$2.71
GPT 5.6 Sol0%0.7327.5s$1.84
Grok 4.50%0.6618.6s$0.61
Floor tests failed here (deduped across repeats)
  • Gemini 3.1 Pro Preview fails Single-influence vertices tear at joints (1/1 runs)
  • GPT 5.6 Sol fails Single-influence vertices tear at joints (1/1 runs)
  • Grok 4.5 fails Single-influence vertices tear at joints (1/1 runs)
coordinate spaces -> run Grok 4.5(cheapest that clears the bar)

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Sonnet 5100%0.239.2s$0.69you are here
GPT 5.6 Sol100%0.2715.4s$1.69
Grok 4.5100%0.4319.1s$0.63run this
Gemini 3.1 Pro Preview67%0.6020.0s$1.92
Floor tests failed here (deduped across repeats)
  • Gemini 3.1 Pro Preview fails Protect a constraint-driven bone from accidental posing (1/1 runs)

Run cost and tokens

Total run spend ~$4.32: generation $1.81 + grading ~$2.51 (judge claude-sonnet-5, estimated). 700,656 tokens = 179,560 generation + 521,096 grading.

modelrequestspromptcompletionreasoningtotal tokenscost
google:gemini-3.1-pro-preview302,22021,67942,75166,650$0.78
openai:gpt-5.6-sol302,24323,94918,35126,578$0.49
anthropic:messages:claude-sonnet-5303,31431,7273,13335,041$0.32
xai:grok-4.53016,41111,15122,26751,291$0.22
generation subtotal12024,18888,50686,502179,560$1.81
grading (judge claude-sonnet-5)442,44678,650-521,096~$2.51 est
total700,656~$4.32

Generation cost is promptfoo's own per-model figure. Grading cost is a LIST-PRICE UPPER BOUND: the judge's grading tokens at its list rate, with no caching or volume discounts, and promptfoo does not meter grading itself. Actual billed cost is usually lower, and provider usage dashboards lag (Anthropic more than most), so confirm real spend in each provider's usage console. Local models are free. The grading estimate assumes ONE constant judge for the whole file (read from the results file's stored config) - if this file was assembled by merging multiple runs made under different judges (e.g. the judge was changed partway through a suite's lifetime and old + new records were combined), the estimate silently uses whichever judge the file currently reports and can misprice the other records' grading tokens.

Suite health

Floor tests failed by many models (likely the TEST, not the models - too strict or naked-recall; review these first)
Floor failures by model
Gemini 3.1 Pro Preview 5 floor tests failed
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
Grok 4.5 5 floor tests failed
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
Claude Sonnet 5 4 floor tests failed
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
GPT 5.6 Sol 4 floor tests failed
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
Saturated discriminating tests (no ranking signal; harden or retire)

Bar: floor pass-rate at or above 100%. Generated by clawhound from a promptfoo results file.