A stronger teacher, a small unconfirmed gain
Learning from deeper search improved some scores, but the Hard result was too uncertain to promote.
· Leaf Arcade experiment notes
The question
Could better policy targets improve move choices without changing the accepted model’s value predictions or tactical features? We collected 768 strategic teaching positions using bounded deeper search, with all actions fully proved in 530 of them.
The method
Two variants used different policy temperatures: one concentrated strongly on the best action, the other spread teaching probability more broadly. Training updated the policy projection and head for 30 epochs, leaving the shared trunk, value head and tactical coefficients fixed.
The softer candidate was frozen before the fresh confirmation. Parent and candidate then faced the same starts at one million exact-search nodes plus 256 PUCT simulations. Both passed 800 out of 800 immediate win/block cases.
The result and decision
Hard rose from 75% to 77% on this bank, but the paired interval included a loss. Better scores on some references did not override the declared Hard promotion rule. We recorded the candidate as an interesting failed experiment and kept the accepted network.
The lesson: a promising point estimate is a reason to investigate a future hypothesis, not a reason to claim confirmed superiority.