Lower value loss did not make a stronger player

A 15.9% improvement in value error did not survive the playing-strength comparison.

· Leaf Arcade experiment notes

What we tried

We trained shared features using stronger teaching data and fresh strategic roots. The hope was that better internal position estimates would improve search planning, particularly against the strongest reference.

The misleading success

Measured value mean-squared error fell from about 0.516 to 0.434, a 15.9% reduction. Position outcome classification barely moved, from 71.9% to 72.2%.

The cheaper candidate did not pass its predeclared playing-strength comparison. Against Hard it scored 71%, versus 70.5% for the full-budget champion on that bank, but the unchanged champion at the same reduced budget scored 72.5%. Training loss alone would have selected the wrong improvement.

What to learn

An auxiliary metric measures one part of a system. A value model, policy and controller interact. The final decision should follow the task metric—playing strength at a fixed runtime budget—while using loss to diagnose what happened.