Run Gemma 4 12b for everyday work: Blender 5.2 API, IK solver behavior. Reach for Claude Haiku 4-5 on coordinate spaces. Reach for Gemini 3.6 Flash on skinning and deformation. No model clears the bar yet on image-to-3D pipeline; review the checks. measurement design has no floor tests yet, so the bar can't be evaluated there - a suite coverage gap, not a model failure.
Headline follows your cheapest lens. The two frontier charts and the per-service table below show both cost and latency.
Cheapest and fastest picks differ on: Blender 5.2 API (cheapest Gemma 4 12b, fastest Claude Sonnet 5); IK solver behavior (cheapest Gemma 4 12b, fastest Claude Haiku 4-5); skinning and deformation (cheapest Gemini 3.6 Flash, fastest Claude Sonnet 5).
Blender 5.2 API-> cheapest Gemma 4 12b, fastest Claude Sonnet 5(the two lenses disagree here)
IK solver behavior-> cheapest Gemma 4 12b, fastest Claude Haiku 4-5(the two lenses disagree here)
image-to-3D pipeline->no model clears it yet(no model clears the bar yet, google:gemini-3.6-flash is closest, review the checks)
measurement design->no floor tests here(this category has no floor tests, so the bar can't be evaluated here; by discriminating score, anthropic:messages:claude-fable-5 ranks highest)
skinning and deformation-> cheapest Gemini 3.6 Flash, fastest Claude Sonnet 5(the two lenses disagree here)
coordinate spaces-> run Claude Haiku 4-5(cheapest that clears the bar)
Cost and latency vs quality
cost vs quality
latency vs quality
Claude Opus 5Gemini 3.6 FlashClaude Fable 5Claude Sonnet 5(you are here)GPT 5.6 SolGPT 5.6 TerraGPT 6 AstraGemma 4 12b+7 more models
All models at a glance
floor
must-pass correctness and guardrails, the routing gate (you want 100%).
disc
discriminating, the graded hard-reasoning score, 0 to 1, used to rank models.
latency
median response time.
cost /100
API cost per 100 tests measured on this suite - it reflects each model's own verbosity here, not just its list price, so it moves if your real tasks are longer (per-test is fractions of a cent; the run-cost panel shows real totals).
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
model
floor
disc
latency
cost /100
Claude Opus 5
93%
0.87
45.0s
$8.33
Gemini 3.6 Flash
79%
0.79
17.6s
$0.76
Claude Fable 5
71%
0.85
23.3s
$8.27
Claude Sonnet 5
71%
0.77
13.1s
$1.08
you are here
Gemma 4 12b*
71%
0.53
24.5s
$0 (free)
GPT 5.6 Sol
71%
0.75
16.8s
$1.68
GPT 5.6 Terra
71%
0.73
11.4s
$0.87
GPT 6 Astra
71%
0.72
20.4s
$4.16
Gemini 3.1 Pro Preview
64%
0.82
25.0s
$2.59
Grok 4.5
64%
0.68
20.6s
$0.76
Qwen 3.8 27b‡
57%
0.71
283.5s
$0 (free)
Grok 4.6
57%
0.69
37.3s
$1.70
Grok 4.3
50%
0.62
9.9s
$0.24
Claude Haiku 4-5
43%
0.54
5.2s
$0.20
DeepSeek R1 32b‡
29%
0.26
160.5s
$0 (free)
Per-service routing
service
cheapest
fastest
Blender 5.2 API
Gemma 4 12b$0 (free)
Claude Sonnet 526.9s
IK solver behavior
Gemma 4 12b$0 (free)
Claude Haiku 4-56.7s
image-to-3D pipeline
no model clears the bar yet (Gemini 3.6 Flash is closest)
measurement design
no floor tests in this category - by discriminating score, Claude Fable 5 ranks highest
skinning and deformation
Gemini 3.6 Flash$0.91/100
Claude Sonnet 519.7s
coordinate spaces
Claude Haiku 4-5cheapest and fastest$0.14/100, 4.1s
Per-category detail
Blender 5.2 API-> run Gemma 4 12b(cheapest that clears the bar)
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
model
floor
disc
latency
cost /100
Claude Sonnet 5
100%
n/a
26.9s
$2.38
you are here
Claude Opus 5
100%
n/a
56.2s
$9.92
Claude Fable 5
100%
n/a
38.1s
$13.33
Gemini 3.1 Pro Preview
100%
n/a
43.5s
$4.63
Gemma 4 12b
100%
n/a
42.2s
$0 (free)
run this
GPT 5.6 Terra
100%
n/a
27.9s
$1.50
GPT 5.6 Sol
100%
n/a
34.9s
$2.42
Grok 4.5
100%
n/a
27.3s
$0.95
GPT 6 Astra
100%
n/a
47.6s
$7.82
Claude Haiku 4-5
50%
n/a
4.8s
$0.18
Gemini 3.6 Flash
50%
n/a
23.5s
$0.79
Qwen 3.8 27b
50%
n/a
n/a
n/a
Grok 4.3
50%
n/a
9.5s
$0.27
Grok 4.6‡
50%
n/a
799.6s
$0.38
DeepSeek R1 32b‡
0%
n/a
113.6s
$0 (free)
Floor tests failed here (deduped across repeats)
Claude Haiku 4-5 fails Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
Gemini 3.6 Flash fails Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
DeepSeek R1 32b fails Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
DeepSeek R1 32b fails Removed Blender 5.2 APIs in an old addon (given the removal facts) (1/1 runs)
Qwen 3.8 27b fails Removed Blender 5.2 APIs in an old addon (given the removal facts) (1/1 runs)
Grok 4.3 fails Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
Grok 4.6 fails Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
IK solver behavior-> run Gemma 4 12b(cheapest that clears the bar)
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
model
floor
disc
latency
cost /100
Claude Haiku 4-5
100%
0.32
6.7s
$0.23
Claude Sonnet 5
100%
0.83
13.4s
$1.02
you are here
Claude Fable 5
100%
0.72
31.5s
$11.44
Claude Opus 5
100%
0.91
60.0s
$11.53
Gemini 3.6 Flash
100%
0.82
18.7s
$0.89
Gemini 3.1 Pro Preview
100%
0.86
28.9s
$2.38
Qwen 3.8 27b‡
100%
0.57
740.4s
$0 (free)
Gemma 4 12b
100%
0.32
30.8s
$0 (free)
run this
GPT 5.6 Terra
100%
0.52
18.7s
$1.01
GPT 5.6 Sol
100%
0.84
29.5s
$2.43
Grok 4.3
100%
0.60
9.2s
$0.27
GPT 6 Astra
100%
0.59
36.9s
$6.68
Grok 4.5
100%
0.73
37.7s
$0.96
Grok 4.6
100%
0.60
70.6s
$2.65
DeepSeek R1 32b‡
50%
0.06
153.0s
$0 (free)
Floor tests failed here (deduped across repeats)
DeepSeek R1 32b fails Pole target distance does not affect the solve (1/1 runs)
image-to-3D pipeline->NO MODEL CLEARS THE BAR - no model clears the bar yet, google:gemini-3.6-flash is closest, review the checks
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
model
floor
disc
latency
cost /100
Gemini 3.6 Flash
83%
n/a
17.8s
$0.65
Claude Opus 5
83%
n/a
44.4s
$7.33
Claude Fable 5
67%
n/a
21.0s
$7.01
Gemma 4 12b
67%
n/a
22.3s
$0 (free)
GPT 5.6 Terra
67%
n/a
9.3s
$0.68
GPT 6 Astra
67%
n/a
20.1s
$3.13
Gemini 3.1 Pro Preview
50%
n/a
21.0s
$2.24
GPT 5.6 Sol
50%
n/a
14.0s
$1.19
Grok 4.6
50%
n/a
44.2s
$1.53
Claude Sonnet 5
33%
n/a
12.4s
$1.02
you are here
Qwen 3.8 27b‡
33%
n/a
186.2s
$0 (free)
Grok 4.5
33%
n/a
22.4s
$0.66
DeepSeek R1 32b‡
17%
n/a
153.0s
$0 (free)
Grok 4.3
17%
n/a
10.7s
$0.26
Claude Haiku 4-5
0%
n/a
5.3s
$0.18
Floor tests failed here (deduped across repeats)
Claude Fable 5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
Claude Fable 5 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Claude Haiku 4-5 fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
Claude Haiku 4-5 fails An asymmetric detail flips sides between the two views (1/1 runs)
Claude Haiku 4-5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
Claude Haiku 4-5 fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
Claude Haiku 4-5 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Claude Haiku 4-5 fails Two-up reference sheet becomes two characters (1/1 runs)
Claude Opus 5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
Claude Sonnet 5 fails An asymmetric detail flips sides between the two views (1/1 runs)
Claude Sonnet 5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
Claude Sonnet 5 fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
Claude Sonnet 5 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Gemini 3.1 Pro Preview fails An asymmetric detail flips sides between the two views (1/1 runs)
Gemini 3.1 Pro Preview fails De-light and PBR, so painted lighting does not double (1/1 runs)
Gemini 3.1 Pro Preview fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
Gemini 3.6 Flash fails De-light and PBR, so painted lighting does not double (1/1 runs)
DeepSeek R1 32b fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
DeepSeek R1 32b fails An asymmetric detail flips sides between the two views (1/1 runs)
DeepSeek R1 32b fails De-light and PBR, so painted lighting does not double (1/1 runs)
DeepSeek R1 32b fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Gemma 4 12b fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
Gemma 4 12b fails De-light and PBR, so painted lighting does not double (1/1 runs)
Qwen 3.8 27b fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
Qwen 3.8 27b fails An asymmetric detail flips sides between the two views (1/1 runs)
Qwen 3.8 27b fails De-light and PBR, so painted lighting does not double (1/1 runs)
Qwen 3.8 27b fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
GPT 5.6 Sol fails An asymmetric detail flips sides between the two views (1/1 runs)
GPT 5.6 Sol fails De-light and PBR, so painted lighting does not double (1/1 runs)
GPT 5.6 Sol fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
GPT 5.6 Terra fails An asymmetric detail flips sides between the two views (1/1 runs)
GPT 5.6 Terra fails De-light and PBR, so painted lighting does not double (1/1 runs)
GPT 6 Astra fails De-light and PBR, so painted lighting does not double (1/1 runs)
GPT 6 Astra fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Grok 4.3 fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
Grok 4.3 fails An asymmetric detail flips sides between the two views (1/1 runs)
Grok 4.3 fails De-light and PBR, so painted lighting does not double (1/1 runs)
Grok 4.3 fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
Grok 4.3 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Grok 4.5 fails An asymmetric detail flips sides between the two views (1/1 runs)
Grok 4.5 fails De-light and PBR, so painted lighting does not double (1/1 runs)
Grok 4.5 fails Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
Grok 4.5 fails Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Grok 4.6 fails A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
Grok 4.6 fails An asymmetric detail flips sides between the two views (1/1 runs)
Grok 4.6 fails De-light and PBR, so painted lighting does not double (1/1 runs)
measurement design->NO FLOOR TESTS HERE - this category has no floor tests, so the bar can't be evaluated here; by discriminating score, anthropic:messages:claude-fable-5 ranks highest
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
model
floor
disc
latency
cost /100
Claude Haiku 4-5
n/a
0.64
5.2s
$0.20
Claude Sonnet 5
n/a
0.83
11.2s
$0.91
you are here
Claude Fable 5
n/a
0.89
23.3s
$7.09
Claude Opus 5
n/a
0.85
41.4s
$7.40
Gemini 3.6 Flash
n/a
0.79
14.1s
$0.76
Gemini 3.1 Pro Preview
n/a
0.84
28.1s
$2.73
Qwen 3.8 27b
n/a
0.76
n/a
n/a
DeepSeek R1 32b‡
n/a
0.30
183.7s
$0 (free)
Gemma 4 12b*
n/a
0.59
20.9s
$0 (free)
GPT 5.6 Terra
n/a
0.74
10.2s
$0.63
GPT 5.6 Sol
n/a
0.78
16.0s
$1.32
Grok 4.3
n/a
0.66
7.4s
$0.20
Grok 4.6
n/a
0.73
32.0s
$1.52
GPT 6 Astra
n/a
0.78
18.7s
$3.08
Grok 4.5
n/a
0.70
19.6s
$0.80
Every model cleared every floor test in this category.
skinning and deformation-> run Gemini 3.6 Flash(cheapest that clears the bar)
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
model
floor
disc
latency
cost /100
Claude Sonnet 5
100%
0.73
19.7s
$1.38
you are here
Claude Fable 5
100%
0.93
26.7s
$8.70
Gemini 3.6 Flash
100%
0.80
21.1s
$0.91
run this
Claude Opus 5
100%
0.94
64.7s
$10.83
Qwen 3.8 27b‡
100%
0.79
283.5s
$0 (free)
Claude Haiku 4-5
0%
0.60
6.6s
$0.25
Gemini 3.1 Pro Preview
0%
0.82
26.0s
$2.71
DeepSeek R1 32b‡
0%
0.34
168.9s
$0 (free)
Gemma 4 12b
0%
0.66
22.1s
$0 (free)
GPT 5.6 Terra
0%
0.86
13.8s
$0.82
GPT 5.6 Sol
0%
0.73
27.5s
$1.84
GPT 6 Astra
0%
0.73
32.4s
$3.83
Grok 4.3
0%
0.72
10.7s
$0.25
Grok 4.5
0%
0.66
18.6s
$0.61
Grok 4.6
0%
0.71
38.0s
$1.19
Floor tests failed here (deduped across repeats)
Claude Haiku 4-5 fails Single-influence vertices tear at joints (1/1 runs)
Gemini 3.1 Pro Preview fails Single-influence vertices tear at joints (1/1 runs)
Grok 4.3 fails Single-influence vertices tear at joints (1/1 runs)
Grok 4.5 fails Single-influence vertices tear at joints (1/1 runs)
Grok 4.6 fails Single-influence vertices tear at joints (1/1 runs)
coordinate spaces-> run Claude Haiku 4-5(cheapest that clears the bar)
Click a column to sort (best first; click again to flip). Hover a row to highlight that model in the charts.
model
floor
disc
latency
cost /100
Claude Haiku 4-5
100%
0.10
4.1s
$0.14
run this
Claude Sonnet 5
100%
0.23
9.2s
$0.69
you are here
Claude Opus 5
100%
0.67
31.5s
$4.61
GPT 5.6 Sol
100%
0.27
15.4s
$1.69
Grok 4.5
100%
0.43
19.1s
$0.63
Grok 4.3
100%
0.13
7.6s
$0.21
Gemini 3.1 Pro Preview
67%
0.60
20.0s
$1.92
Gemini 3.6 Flash
67%
0.63
16.0s
$0.63
Qwen 3.8 27b
67%
0.37
n/a
n/a
DeepSeek R1 32b‡
67%
0.30
111.0s
$0 (free)
Gemma 4 12b
67%
0.27
32.3s
$0 (free)
GPT 5.6 Terra
67%
0.90
10.5s
$1.15
GPT 6 Astra
67%
0.53
14.1s
$3.24
Grok 4.6
67%
0.60
59.3s
$2.34
Claude Fable 5
33%
0.67
21.5s
$5.86
Floor tests failed here (deduped across repeats)
Claude Fable 5 fails Armature-local to world coordinate conversion (given the mapping) (1/1 runs)
Claude Fable 5 fails Protect a constraint-driven bone from accidental posing (1/1 runs)
Gemini 3.1 Pro Preview fails Protect a constraint-driven bone from accidental posing (1/1 runs)
Gemini 3.6 Flash fails Protect a constraint-driven bone from accidental posing (1/1 runs)
DeepSeek R1 32b fails Protect a constraint-driven bone from accidental posing (1/1 runs)
Gemma 4 12b fails Protect a constraint-driven bone from accidental posing (1/1 runs)
Qwen 3.8 27b fails Protect a constraint-driven bone from accidental posing (1/1 runs)
GPT 5.6 Terra fails Protect a constraint-driven bone from accidental posing (1/1 runs)
GPT 6 Astra fails Armature-local to world coordinate conversion (given the mapping) (1/1 runs)
Grok 4.6 fails Protect a constraint-driven bone from accidental posing (1/1 runs)
Run cost and tokens
Total run spend ~$18.75: generation $9.10 + grading ~$9.65 (judge claude-sonnet-5, estimated). 2,768,232 tokens = 788,356 generation + 1,979,876 grading.
model
requests
prompt
completion
reasoning
total tokens
cost
anthropic:messages:claude-opus-5
30
3,314
99,270
57,062
102,584
$2.50
anthropic:messages:claude-fable-5
30
3,314
48,932
23,896
52,246
$2.48
openai:gpt-6-astra
30
2,243
23,696
16,917
26,487
$1.21
google:gemini-3.1-pro-preview
30
2,220
21,679
42,751
66,650
$0.78
xai:grok-4.6
30
19,816
7,139
71,864
101,427
$0.49
openai:gpt-5.6-sol
30
2,243
23,949
18,351
26,578
$0.49
anthropic:messages:claude-sonnet-5
30
3,314
31,727
3,133
35,041
$0.32
openai:gpt-5.6-terra
30
2,243
20,586
11,187
23,197
$0.25
google:gemini-3.6-flash
30
2,220
20,339
40,268
62,827
$0.23
xai:grok-4.5
30
16,411
11,151
22,267
51,291
$0.22
xai:grok-4.3
30
7,608
6,425
19,080
34,318
$0.07
anthropic:messages:claude-haiku-4-5
30
2,510
11,312
0
13,822
$0.06
ollama:chat:qwen3.8:27b-q4_K_M
30
2,452
79,808
0
82,260
$0.00 (local)
ollama:chat:gemma4:12b-it-q4_K_M
30
2,670
71,629
0
74,299
$0.00 (local)
ollama:chat:deepseek-r1:32b
30
2,244
33,085
0
35,329
$0.00 (local)
generation subtotal
450
74,822
510,727
326,776
788,356
$9.10
grading (judge claude-sonnet-5)
1,670,456
309,420
-
1,979,876
~$9.65 est
total
2,768,232
~$18.75
Generation cost is promptfoo's own per-model figure. Grading cost is a LIST-PRICE UPPER BOUND: the judge's grading tokens at its list rate, with no caching or volume discounts, and promptfoo does not meter grading itself. Actual billed cost is usually lower, and provider usage dashboards lag (Anthropic more than most), so confirm real spend in each provider's usage console. Local models are free. The grading estimate assumes ONE constant judge for the whole file (read from the results file's stored config) - if this file was assembled by merging multiple runs made under different judges (e.g. the judge was changed partway through a suite's lifetime and old + new records were combined), the estimate silently uses whichever judge the file currently reports and can misprice the other records' grading tokens.
Suite health
Floor tests failed by many models (likely the TEST, not the models - too strict or naked-recall; review these first)
De-light and PBR, so painted lighting does not double - fails 15/15 models
An asymmetric detail flips sides between the two views - fails 10/15 models
Single-influence vertices tear at joints - fails 10/15 models
Palms-down T-pose hides the thumbs from reconstruction - fails 9/15 models
Protect a constraint-driven bone from accidental posing - fails 8/15 models
Floor failures by model
DeepSeek R1 32b10 floor tests failed
A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
An asymmetric detail flips sides between the two views (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Pole target distance does not affect the solve (1/1 runs)
Protect a constraint-driven bone from accidental posing (1/1 runs)
Removed Blender 5.2 APIs in an old addon (given the removal facts) (1/1 runs)
Single-influence vertices tear at joints (1/1 runs)
Two-up reference sheet becomes two characters (1/1 runs)
Claude Haiku 4-58 floor tests failed
A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
An asymmetric detail flips sides between the two views (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Single-influence vertices tear at joints (1/1 runs)
Two-up reference sheet becomes two characters (1/1 runs)
Grok 4.37 floor tests failed
A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
An asymmetric detail flips sides between the two views (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Single-influence vertices tear at joints (1/1 runs)
Qwen 3.8 27b6 floor tests failed
A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
An asymmetric detail flips sides between the two views (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Protect a constraint-driven bone from accidental posing (1/1 runs)
Removed Blender 5.2 APIs in an old addon (given the removal facts) (1/1 runs)
Grok 4.66 floor tests failed
A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
An asymmetric detail flips sides between the two views (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
Protect a constraint-driven bone from accidental posing (1/1 runs)
Single-influence vertices tear at joints (1/1 runs)
Gemini 3.1 Pro Preview5 floor tests failed
An asymmetric detail flips sides between the two views (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
Protect a constraint-driven bone from accidental posing (1/1 runs)
Single-influence vertices tear at joints (1/1 runs)
Grok 4.55 floor tests failed
An asymmetric detail flips sides between the two views (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Single-influence vertices tear at joints (1/1 runs)
Claude Fable 54 floor tests failed
Armature-local to world coordinate conversion (given the mapping) (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Protect a constraint-driven bone from accidental posing (1/1 runs)
Claude Sonnet 54 floor tests failed
An asymmetric detail flips sides between the two views (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Mixamo misreads an unrigged FBX as already-rigged (1/1 runs)
Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Gemma 4 12b4 floor tests failed
A transparent PNG background flattens to black and kills the silhouette (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Protect a constraint-driven bone from accidental posing (1/1 runs)
Single-influence vertices tear at joints (1/1 runs)
GPT 5.6 Sol4 floor tests failed
An asymmetric detail flips sides between the two views (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Single-influence vertices tear at joints (1/1 runs)
GPT 5.6 Terra4 floor tests failed
An asymmetric detail flips sides between the two views (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Protect a constraint-driven bone from accidental posing (1/1 runs)
Single-influence vertices tear at joints (1/1 runs)
GPT 6 Astra4 floor tests failed
Armature-local to world coordinate conversion (given the mapping) (1/1 runs)
De-light and PBR, so painted lighting does not double (1/1 runs)
Palms-down T-pose hides the thumbs from reconstruction (1/1 runs)
Single-influence vertices tear at joints (1/1 runs)
Gemini 3.6 Flash3 floor tests failed
De-light and PBR, so painted lighting does not double (1/1 runs)
Euler order to isolate twist from bend (given the contamination rule) (1/1 runs)
Protect a constraint-driven bone from accidental posing (1/1 runs)
Claude Opus 51 floor test failed
De-light and PBR, so painted lighting does not double (1/1 runs)
Bar: floor pass-rate at or above 100%. Generated by clawhound from a promptfoo results file.