Engineering Bench · CAD track · public leaderboard
Who grades the CAD models? A physics solver, not another LLM.
Six rungs — compose, holes, survive, diagnose+fix, optimize, assembly. Geometry verified
(Chamfer < 0.5% of the part diagonal) plus a CalculiX FEA solve (safety factor ≥ 1.5); tasks
regenerate from seed, so there is nothing to memorize. No model passes rung 5 (optimize) or
rung 6 (assembly).
Engineering Bench — CAD trackloading…
Rung 5–6 · unconqueredreceipts for every run on this board
Graded on the OUTCOME (STEP geometry + FEA), never the method — the model may use tools or write code, its choice.
Why physics is the trust
The grader is a solver, not a scorer. Every output part is meshed and solved with CalculiX against steel (yield 250 MPa): von Mises, safety factor, displacement, strain.
Reproducible = auditable. Same task, same solver, same result — no judge-model drift, no prompt-variance noise.
Match is measured too. Chamfer distance < 0.5% of the part diagonal, so a model can't pass physics on geometry that isn't the part.
Anti-contamination. Dimensions and loads regenerate from the seed every run — the answer key can't leak into training data.
The six rungs
1 compose — build a part from primitives
2 holes — feature placement
3 survive — strength under load
4 diagnose+fix — find and repair a failing design
5 optimize — improve a real objective
6 assembly — multi-part reasoning
Get your model graded
You built a CAD-generating model (or a CAD agent) and want a number that survives scrutiny? Every engagement starts with the failure map — you see where your model breaks before you see an invoice.
Free taste
$0
10 engineering tasks, one domain, publishable score + failure map. 2 days. First serious team only.
Sprint
$5K
100 tasks across all six rungs, full failure map with FEA evidence, publishable score. 5 days.
Deep
$10K
Your domain + 300 tasks + a verified bad/good data pack with FEA labels. 1–2 weeks.