Run Grok 4.5 for everyday work: Blender 5.2 API, IK solver behavior, coordinate spaces. Reach for Claude Sonnet 5 on skinning and deformation. No model clears the bar yet on image-to-3D pipeline; review the checks. measurement design has no floor tests yet, so the bar can't be evaluated there - a suite coverage gap, not a model failure.
Headline follows your cheapest lens. The two frontier charts and the per-service table below show both cost and latency.
Cheapest and fastest picks differ on: Blender 5.2 API (cheapest Grok 4.5, fastest Claude Sonnet 5); IK solver behavior (cheapest Grok 4.5, fastest Claude Sonnet 5); coordinate spaces (cheapest Grok 4.5, fastest Claude Sonnet 5).
Benchmark: Claude Sonnet 5. Positive cost/latency is worse (more expensive/slower than the benchmark); positive disc is better (higher quality).
| model | cost vs benchmark | latency vs benchmark | disc vs benchmark |
|---|---|---|---|
| Grok 4.5 | -30% | +58% | -0.09 |
| GPT 5.6 Sol | +56% | +28% | -0.02 |
| Gemini 3.1 Pro Preview | +140% | +91% | +0.05 |
| floor | must-pass correctness and guardrails, the routing gate (you want 100%). |
| disc | discriminating, the graded hard-reasoning score, 0 to 1, used to rank models. |
| latency | median response time. |
| cost /100 | API cost per 100 tests measured on this suite - it reflects each model's own verbosity here, not just its list price, so it moves if your real tasks are longer (per-test is fractions of a cent; the run-cost panel shows real totals). |
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
| model | floor | disc | latency | cost /100 | |
|---|---|---|---|---|---|
| Claude Sonnet 5 | 71% | 0.77 | 13.1s | $1.08 | you are here |
| GPT 5.6 Sol | 71% | 0.75 | 16.8s | $1.68 | |
| Gemini 3.1 Pro Preview | 64% | 0.82 | 25.0s | $2.59 | |
| Grok 4.5 | 64% | 0.68 | 20.6s | $0.76 |
| service | cheapest | fastest |
|---|---|---|
| Blender 5.2 API | Grok 4.5 $0.95/100 | Claude Sonnet 5 26.9s |
| IK solver behavior | Grok 4.5 $0.96/100 | Claude Sonnet 5 13.4s |
| image-to-3D pipeline | no model clears the bar yet (GPT 5.6 Sol is closest) | |
| measurement design | no floor tests in this category - by discriminating score, Gemini 3.1 Pro Preview ranks highest | |
| skinning and deformation | Claude Sonnet 5 cheapest and fastest $1.38/100, 19.7s | |
| coordinate spaces | Grok 4.5 $0.63/100 | Claude Sonnet 5 9.2s |
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
| model | floor | disc | latency | cost /100 | |
|---|---|---|---|---|---|
| Claude Sonnet 5 | 100% | n/a | 26.9s | $2.38 | you are here |
| Gemini 3.1 Pro Preview | 100% | n/a | 43.5s | $4.63 | |
| GPT 5.6 Sol | 100% | n/a | 34.9s | $2.42 | |
| Grok 4.5 | 100% | n/a | 27.3s | $0.95 | run this |
Every model cleared every floor test in this category.
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
| model | floor | disc | latency | cost /100 | |
|---|---|---|---|---|---|
| Claude Sonnet 5 | 100% | 0.83 | 13.4s | $1.02 | you are here |
| Gemini 3.1 Pro Preview | 100% | 0.86 | 28.9s | $2.38 | |
| GPT 5.6 Sol | 100% | 0.84 | 29.5s | $2.43 | |
| Grok 4.5 | 100% | 0.73 | 37.7s | $0.96 | run this |
Every model cleared every floor test in this category.
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
| model | floor | disc | latency | cost /100 | |
|---|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 50% | n/a | 21.0s | $2.24 | |
| GPT 5.6 Sol | 50% | n/a | 14.0s | $1.19 | |
| Claude Sonnet 5 | 33% | n/a | 12.4s | $1.02 | you are here |
| Grok 4.5 | 33% | n/a | 22.4s | $0.66 |
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
| model | floor | disc | latency | cost /100 | |
|---|---|---|---|---|---|
| Claude Sonnet 5 | n/a | 0.83 | 11.2s | $0.91 | you are here |
| Gemini 3.1 Pro Preview | n/a | 0.84 | 28.1s | $2.73 | |
| GPT 5.6 Sol | n/a | 0.78 | 16.0s | $1.32 | |
| Grok 4.5 | n/a | 0.70 | 19.6s | $0.80 |
Every model cleared every floor test in this category.
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
| model | floor | disc | latency | cost /100 | |
|---|---|---|---|---|---|
| Claude Sonnet 5 | 100% | 0.73 | 19.7s | $1.38 | run this |
| Gemini 3.1 Pro Preview | 0% | 0.82 | 26.0s | $2.71 | |
| GPT 5.6 Sol | 0% | 0.73 | 27.5s | $1.84 | |
| Grok 4.5 | 0% | 0.66 | 18.6s | $0.61 |
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
| model | floor | disc | latency | cost /100 | |
|---|---|---|---|---|---|
| Claude Sonnet 5 | 100% | 0.23 | 9.2s | $0.69 | you are here |
| GPT 5.6 Sol | 100% | 0.27 | 15.4s | $1.69 | |
| Grok 4.5 | 100% | 0.43 | 19.1s | $0.63 | run this |
| Gemini 3.1 Pro Preview | 67% | 0.60 | 20.0s | $1.92 |
Total run spend ~$4.32: generation $1.81 + grading ~$2.51 (judge claude-sonnet-5, estimated). 700,656 tokens = 179,560 generation + 521,096 grading.
| model | requests | prompt | completion | reasoning | total tokens | cost |
|---|---|---|---|---|---|---|
| google:gemini-3.1-pro-preview | 30 | 2,220 | 21,679 | 42,751 | 66,650 | $0.78 |
| openai:gpt-5.6-sol | 30 | 2,243 | 23,949 | 18,351 | 26,578 | $0.49 |
| anthropic:messages:claude-sonnet-5 | 30 | 3,314 | 31,727 | 3,133 | 35,041 | $0.32 |
| xai:grok-4.5 | 30 | 16,411 | 11,151 | 22,267 | 51,291 | $0.22 |
| generation subtotal | 120 | 24,188 | 88,506 | 86,502 | 179,560 | $1.81 |
| grading (judge claude-sonnet-5) | 442,446 | 78,650 | - | 521,096 | ~$2.51 est | |
| total | 700,656 | ~$4.32 |
Generation cost is promptfoo's own per-model figure. Grading cost is a LIST-PRICE UPPER BOUND: the judge's grading tokens at its list rate, with no caching or volume discounts, and promptfoo does not meter grading itself. Actual billed cost is usually lower, and provider usage dashboards lag (Anthropic more than most), so confirm real spend in each provider's usage console. Local models are free. The grading estimate assumes ONE constant judge for the whole file (read from the results file's stored config) - if this file was assembled by merging multiple runs made under different judges (e.g. the judge was changed partway through a suite's lifetime and old + new records were combined), the estimate silently uses whichever judge the file currently reports and can misprice the other records' grading tokens.
Bar: floor pass-rate at or above 100%. Generated by clawhound from a promptfoo results file.