Every held-out position, with each model’s recommended move on the same grounded input. Rows where OURS wins— sound and tier-fit where a frontier model isn’t, or faithful where one invents a fact — are surfaced first.
◐Prior-generation benchmark (chess-coach-v2 (1.7B))This 200-position held-out study was run on the earlier 1.7B (v2) coach. The live coach now serves v6-dpo2 (Qwen3-32B); for the curated, per-level comparison — OURS (v4) vs the frontier on genuine tier forks, with the live re-run (now v6-dpo2) — see the Multi-Model Showcase.