CADGenBench Local Analysis — GPT-5.5 / X3 / Gemini 3.5 Flash / Kimi K2.6
Local metric is CADGenBench sanity-check validity (valid + watertight STEP) only. Final CAD Score requires the official leaderboard because ground truth is private. Models use the same protocol: raw pass (generic prompt) + repair pass (constrained build123d prompt + previous-error feedback); best-of picks repair when valid else raw.
Generated at 2026-07-02 02:28:19 · 81 samples
Model summary & differential (best-of vs best-of)
Unified Benchmark Table (local validity + official CAD Score baselines)
Blue rows = this run (local validity only). Other rows = official CADGenBench leaderboard (real CAD Score on private ground truth).