← back to the article

Which model to run

Run Gemma 4 12b for everyday work: Blender 5.2 API, IK solver behavior. Reach for Claude Haiku 4-5 on coordinate spaces. Reach for Gemini 3.6 Flash on skinning and deformation. No model clears the bar yet on image-to-3D pipeline; review the checks. measurement design has no floor tests yet, so the bar can't be evaluated there - a suite coverage gap, not a model failure.

Headline follows your cheapest lens. The two frontier charts and the per-service table below show both cost and latency.

Cheapest and fastest picks differ on: Blender 5.2 API (cheapest Gemma 4 12b, fastest Claude Sonnet 5); IK solver behavior (cheapest Gemma 4 12b, fastest Claude Haiku 4-5); skinning and deformation (cheapest Gemini 3.6 Flash, fastest Claude Sonnet 5).

Cost and latency vs quality

cost vs quality
28%46%64%82%100%$0.00$2.39$4.79$7.18$9.58cost per 100 tests (USD), lower is betterfloor pass-rate+7 more
latency vs quality
28%46%64%82%100%0.000s12.926s25.852s38.778s51.704smedian latency per test (s), lower is betterfloor pass-rate+7 more
Claude Opus 5Gemini 3.6 FlashClaude Fable 5Claude Sonnet 5 (you are here)GPT 5.6 SolGPT 5.6 TerraGPT 6 AstraGemma 4 12b+7 more models

All models at a glance

floormust-pass correctness and guardrails, the routing gate (you want 100%).
discdiscriminating, the graded hard-reasoning score, 0 to 1, used to rank models.
latencymedian response time.
cost /100API cost per 100 tests measured on this suite - it reflects each model's own verbosity here, not just its list price, so it moves if your real tasks are longer (per-test is fractions of a cent; the run-cost panel shows real totals).

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Opus 593%0.8745.0s$8.33
Gemini 3.6 Flash79%0.7917.6s$0.76
Claude Fable 571%0.8523.3s$8.27
Claude Sonnet 571%0.7713.1s$1.08you are here
Gemma 4 12b*71%0.5324.5s$0 (free)
GPT 5.6 Sol71%0.7516.8s$1.68
GPT 5.6 Terra71%0.7311.4s$0.87
GPT 6 Astra71%0.7220.4s$4.16
Gemini 3.1 Pro Preview64%0.8225.0s$2.59
Grok 4.564%0.6820.6s$0.76
Qwen 3.8 27b‡57%0.71283.5s$0 (free)
Grok 4.657%0.6937.3s$1.70
Grok 4.350%0.629.9s$0.24
Claude Haiku 4-543%0.545.2s$0.20
DeepSeek R1 32b‡29%0.26160.5s$0 (free)

Per-service routing

servicecheapestfastest
Blender 5.2 APIGemma 4 12b $0 (free)Claude Sonnet 5 26.9s
IK solver behaviorGemma 4 12b $0 (free)Claude Haiku 4-5 6.7s
image-to-3D pipelineno model clears the bar yet (Gemini 3.6 Flash is closest)
measurement designno floor tests in this category - by discriminating score, Claude Fable 5 ranks highest
skinning and deformationGemini 3.6 Flash $0.91/100Claude Sonnet 5 19.7s
coordinate spacesClaude Haiku 4-5 cheapest and fastest $0.14/100, 4.1s

Per-category detail

Blender 5.2 API -> run Gemma 4 12b(cheapest that clears the bar)

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Sonnet 5100%n/a26.9s$2.38you are here
Claude Opus 5100%n/a56.2s$9.92
Claude Fable 5100%n/a38.1s$13.33
Gemini 3.1 Pro Preview100%n/a43.5s$4.63
Gemma 4 12b100%n/a42.2s$0 (free)run this
GPT 5.6 Terra100%n/a27.9s$1.50
GPT 5.6 Sol100%n/a34.9s$2.42
Grok 4.5100%n/a27.3s$0.95
GPT 6 Astra100%n/a47.6s$7.82
Claude Haiku 4-550%n/a4.8s$0.18
Gemini 3.6 Flash50%n/a23.5s$0.79
Qwen 3.8 27b50%n/an/an/a
Grok 4.350%n/a9.5s$0.27
Grok 4.6‡50%n/a799.6s$0.38
DeepSeek R1 32b‡0%n/a113.6s$0 (free)
Floor tests failed here (deduped across repeats)
  • Claude Haiku 4-5 fails Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
  • Gemini 3.6 Flash fails Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
  • DeepSeek R1 32b fails Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
  • DeepSeek R1 32b fails Removed Blender 5.2 APIs in an old addon (given the removal facts) (1/1 runs)
  • Qwen 3.8 27b fails Removed Blender 5.2 APIs in an old addon (given the removal facts) (1/1 runs)
  • Grok 4.3 fails Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
  • Grok 4.6 fails Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
IK solver behavior -> run Gemma 4 12b(cheapest that clears the bar)

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Haiku 4-5100%0.326.7s$0.23
Claude Sonnet 5100%0.8313.4s$1.02you are here
Claude Fable 5100%0.7231.5s$11.44
Claude Opus 5100%0.9160.0s$11.53
Gemini 3.6 Flash100%0.8218.7s$0.89
Gemini 3.1 Pro Preview100%0.8628.9s$2.38
Qwen 3.8 27b‡100%0.57740.4s$0 (free)
Gemma 4 12b100%0.3230.8s$0 (free)run this
GPT 5.6 Terra100%0.5218.7s$1.01
GPT 5.6 Sol100%0.8429.5s$2.43
Grok 4.3100%0.609.2s$0.27
GPT 6 Astra100%0.5936.9s$6.68
Grok 4.5100%0.7337.7s$0.96
Grok 4.6100%0.6070.6s$2.65
DeepSeek R1 32b‡50%0.06153.0s$0 (free)
Floor tests failed here (deduped across repeats)
  • DeepSeek R1 32b fails Pole target distance does not affect the solve (1/1 runs)
image-to-3D pipeline -> NO MODEL CLEARS THE BAR - no model clears the bar yet, google:gemini-3.6-flash is closest, review the checks

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Gemini 3.6 Flash83%n/a17.8s$0.65
Claude Opus 583%n/a44.4s$7.33
Claude Fable 567%n/a21.0s$7.01
Gemma 4 12b67%n/a22.3s$0 (free)
GPT 5.6 Terra67%n/a9.3s$0.68
GPT 6 Astra67%n/a20.1s$3.13
Gemini 3.1 Pro Preview50%n/a21.0s$2.24
GPT 5.6 Sol50%n/a14.0s$1.19
Grok 4.650%n/a44.2s$1.53
Claude Sonnet 533%n/a12.4s$1.02you are here
Qwen 3.8 27b‡33%n/a186.2s$0 (free)
Grok 4.533%n/a22.4s$0.66
DeepSeek R1 32b‡17%n/a153.0s$0 (free)
Grok 4.317%n/a10.7s$0.26
Claude Haiku 4-50%n/a5.3s$0.18
Floor tests failed here (deduped across repeats)
  • Claude Fable 5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Claude Fable 5 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Claude Haiku 4-5 fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • Claude Haiku 4-5 fails An asymmetric detail flips sides between the two views (1/1 runs)
  • Claude Haiku 4-5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Claude Haiku 4-5 fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Claude Haiku 4-5 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Claude Haiku 4-5 fails Two-up reference sheet becomes two characters (1/1 runs)
  • Claude Opus 5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Claude Sonnet 5 fails An asymmetric detail flips sides between the two views (1/1 runs)
  • Claude Sonnet 5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Claude Sonnet 5 fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Claude Sonnet 5 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Gemini 3.1 Pro Preview fails An asymmetric detail flips sides between the two views (1/1 runs)
  • Gemini 3.1 Pro Preview fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Gemini 3.1 Pro Preview fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Gemini 3.6 Flash fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • DeepSeek R1 32b fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • DeepSeek R1 32b fails An asymmetric detail flips sides between the two views (1/1 runs)
  • DeepSeek R1 32b fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • DeepSeek R1 32b fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • DeepSeek R1 32b fails Two-up reference sheet becomes two characters (1/1 runs)
  • Gemma 4 12b fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • Gemma 4 12b fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Qwen 3.8 27b fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • Qwen 3.8 27b fails An asymmetric detail flips sides between the two views (1/1 runs)
  • Qwen 3.8 27b fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Qwen 3.8 27b fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • GPT 5.6 Sol fails An asymmetric detail flips sides between the two views (1/1 runs)
  • GPT 5.6 Sol fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • GPT 5.6 Sol fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • GPT 5.6 Terra fails An asymmetric detail flips sides between the two views (1/1 runs)
  • GPT 5.6 Terra fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • GPT 6 Astra fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • GPT 6 Astra fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Grok 4.3 fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • Grok 4.3 fails An asymmetric detail flips sides between the two views (1/1 runs)
  • Grok 4.3 fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Grok 4.3 fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Grok 4.3 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Grok 4.5 fails An asymmetric detail flips sides between the two views (1/1 runs)
  • Grok 4.5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
  • Grok 4.5 fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Grok 4.5 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Grok 4.6 fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • Grok 4.6 fails An asymmetric detail flips sides between the two views (1/1 runs)
  • Grok 4.6 fails De-light and PBR, so painted lighting does not double (1/1 runs)
measurement design -> NO FLOOR TESTS HERE - this category has no floor tests, so the bar can't be evaluated here; by discriminating score, anthropic:messages:claude-fable-5 ranks highest

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Haiku 4-5n/a0.645.2s$0.20
Claude Sonnet 5n/a0.8311.2s$0.91you are here
Claude Fable 5n/a0.8923.3s$7.09
Claude Opus 5n/a0.8541.4s$7.40
Gemini 3.6 Flashn/a0.7914.1s$0.76
Gemini 3.1 Pro Previewn/a0.8428.1s$2.73
Qwen 3.8 27bn/a0.76n/an/a
DeepSeek R1 32b‡n/a0.30183.7s$0 (free)
Gemma 4 12b*n/a0.5920.9s$0 (free)
GPT 5.6 Terran/a0.7410.2s$0.63
GPT 5.6 Soln/a0.7816.0s$1.32
Grok 4.3n/a0.667.4s$0.20
Grok 4.6n/a0.7332.0s$1.52
GPT 6 Astran/a0.7818.7s$3.08
Grok 4.5n/a0.7019.6s$0.80

Every model cleared every floor test in this category.

skinning and deformation -> run Gemini 3.6 Flash(cheapest that clears the bar)

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Sonnet 5100%0.7319.7s$1.38you are here
Claude Fable 5100%0.9326.7s$8.70
Gemini 3.6 Flash100%0.8021.1s$0.91run this
Claude Opus 5100%0.9464.7s$10.83
Qwen 3.8 27b‡100%0.79283.5s$0 (free)
Claude Haiku 4-50%0.606.6s$0.25
Gemini 3.1 Pro Preview0%0.8226.0s$2.71
DeepSeek R1 32b‡0%0.34168.9s$0 (free)
Gemma 4 12b0%0.6622.1s$0 (free)
GPT 5.6 Terra0%0.8613.8s$0.82
GPT 5.6 Sol0%0.7327.5s$1.84
GPT 6 Astra0%0.7332.4s$3.83
Grok 4.30%0.7210.7s$0.25
Grok 4.50%0.6618.6s$0.61
Grok 4.60%0.7138.0s$1.19
Floor tests failed here (deduped across repeats)
  • Claude Haiku 4-5 fails Single-influence vertices tear at joints (1/1 runs)
  • Gemini 3.1 Pro Preview fails Single-influence vertices tear at joints (1/1 runs)
  • DeepSeek R1 32b fails Single-influence vertices tear at joints (1/1 runs)
  • Gemma 4 12b fails Single-influence vertices tear at joints (1/1 runs)
  • GPT 5.6 Sol fails Single-influence vertices tear at joints (1/1 runs)
  • GPT 5.6 Terra fails Single-influence vertices tear at joints (1/1 runs)
  • GPT 6 Astra fails Single-influence vertices tear at joints (1/1 runs)
  • Grok 4.3 fails Single-influence vertices tear at joints (1/1 runs)
  • Grok 4.5 fails Single-influence vertices tear at joints (1/1 runs)
  • Grok 4.6 fails Single-influence vertices tear at joints (1/1 runs)
coordinate spaces -> run Claude Haiku 4-5(cheapest that clears the bar)

Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.

modelfloordisclatencycost /100
Claude Haiku 4-5100%0.104.1s$0.14run this
Claude Sonnet 5100%0.239.2s$0.69you are here
Claude Opus 5100%0.6731.5s$4.61
GPT 5.6 Sol100%0.2715.4s$1.69
Grok 4.5100%0.4319.1s$0.63
Grok 4.3100%0.137.6s$0.21
Gemini 3.1 Pro Preview67%0.6020.0s$1.92
Gemini 3.6 Flash67%0.6316.0s$0.63
Qwen 3.8 27b67%0.37n/an/a
DeepSeek R1 32b‡67%0.30111.0s$0 (free)
Gemma 4 12b67%0.2732.3s$0 (free)
GPT 5.6 Terra67%0.9010.5s$1.15
GPT 6 Astra67%0.5314.1s$3.24
Grok 4.667%0.6059.3s$2.34
Claude Fable 533%0.6721.5s$5.86
Floor tests failed here (deduped across repeats)
  • Claude Fable 5 fails Armature-local to world coordinate conversion (given the mapping) (1/1 runs)
  • Claude Fable 5 fails Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Gemini 3.1 Pro Preview fails Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Gemini 3.6 Flash fails Protect a constraint-driven bone from accidental posing (1/1 runs)
  • DeepSeek R1 32b fails Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Gemma 4 12b fails Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Qwen 3.8 27b fails Protect a constraint-driven bone from accidental posing (1/1 runs)
  • GPT 5.6 Terra fails Protect a constraint-driven bone from accidental posing (1/1 runs)
  • GPT 6 Astra fails Armature-local to world coordinate conversion (given the mapping) (1/1 runs)
  • Grok 4.6 fails Protect a constraint-driven bone from accidental posing (1/1 runs)

Run cost and tokens

Total run spend ~$18.75: generation $9.10 + grading ~$9.65 (judge claude-sonnet-5, estimated). 2,768,232 tokens = 788,356 generation + 1,979,876 grading.

modelrequestspromptcompletionreasoningtotal tokenscost
anthropic:messages:claude-opus-5303,31499,27057,062102,584$2.50
anthropic:messages:claude-fable-5303,31448,93223,89652,246$2.48
openai:gpt-6-astra302,24323,69616,91726,487$1.21
google:gemini-3.1-pro-preview302,22021,67942,75166,650$0.78
xai:grok-4.63019,8167,13971,864101,427$0.49
openai:gpt-5.6-sol302,24323,94918,35126,578$0.49
anthropic:messages:claude-sonnet-5303,31431,7273,13335,041$0.32
openai:gpt-5.6-terra302,24320,58611,18723,197$0.25
google:gemini-3.6-flash302,22020,33940,26862,827$0.23
xai:grok-4.53016,41111,15122,26751,291$0.22
xai:grok-4.3307,6086,42519,08034,318$0.07
anthropic:messages:claude-haiku-4-5302,51011,312013,822$0.06
ollama:chat:qwen3.8:27b-q4_K_M302,45279,808082,260$0.00 (local)
ollama:chat:gemma4:12b-it-q4_K_M302,67071,629074,299$0.00 (local)
ollama:chat:deepseek-r1:32b302,24433,085035,329$0.00 (local)
generation subtotal45074,822510,727326,776788,356$9.10
grading (judge claude-sonnet-5)1,670,456309,420-1,979,876~$9.65 est
total2,768,232~$18.75

Generation cost is promptfoo's own per-model figure. Grading cost is a LIST-PRICE UPPER BOUND: the judge's grading tokens at its list rate, with no caching or volume discounts, and promptfoo does not meter grading itself. Actual billed cost is usually lower, and provider usage dashboards lag (Anthropic more than most), so confirm real spend in each provider's usage console. Local models are free. The grading estimate assumes ONE constant judge for the whole file (read from the results file's stored config) - if this file was assembled by merging multiple runs made under different judges (e.g. the judge was changed partway through a suite's lifetime and old + new records were combined), the estimate silently uses whichever judge the file currently reports and can misprice the other records' grading tokens.

Suite health

Floor tests failed by many models (likely the TEST, not the models - too strict or naked-recall; review these first)
Floor failures by model
DeepSeek R1 32b 10 floor tests failed
  • A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Pole target distance does not affect the solve (1/1 runs)
  • Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Removed Blender 5.2 APIs in an old addon (given the removal facts) (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
  • Two-up reference sheet becomes two characters (1/1 runs)
Claude Haiku 4-5 8 floor tests failed
  • A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
  • Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
  • Two-up reference sheet becomes two characters (1/1 runs)
Grok 4.3 7 floor tests failed
  • A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
  • Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
Qwen 3.8 27b 6 floor tests failed
  • A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Removed Blender 5.2 APIs in an old addon (given the removal facts) (1/1 runs)
Grok 4.6 6 floor tests failed
  • A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
  • Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
Gemini 3.1 Pro Preview 5 floor tests failed
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
Grok 4.5 5 floor tests failed
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
Claude Fable 5 4 floor tests failed
  • Armature-local to world coordinate conversion (given the mapping) (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Protect a constraint-driven bone from accidental posing (1/1 runs)
Claude Sonnet 5 4 floor tests failed
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Gemma 4 12b 4 floor tests failed
  • A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
GPT 5.6 Sol 4 floor tests failed
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
GPT 5.6 Terra 4 floor tests failed
  • An asymmetric detail flips sides between the two views (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Protect a constraint-driven bone from accidental posing (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
GPT 6 Astra 4 floor tests failed
  • Armature-local to world coordinate conversion (given the mapping) (1/1 runs)
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
  • Single-influence vertices tear at joints (1/1 runs)
Gemini 3.6 Flash 3 floor tests failed
  • De-light and PBR, so painted lighting does not double (1/1 runs)
  • Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
  • Protect a constraint-driven bone from accidental posing (1/1 runs)
Claude Opus 5 1 floor test failed
  • De-light and PBR, so painted lighting does not double (1/1 runs)

Bar: floor pass-rate at or above 100%. Generated by clawhound from a promptfoo results file.