When is the model good enough?
Three experiments failed to establish a useful gain. We kept the champion and moved effort into a usable learning product.
· Leaf Arcade experiment notes
Define a useful improvement before training
A stronger Connect Four model must retain every certified immediate win and block, avoid material regressions, and improve on fresh paired games at the same search budget. For the last round, the rule was a gain of at least three percentage points against Hard and a positive lower 95% paired confidence bound.
A cheaper model needs its own strength and latency comparison. Fewer search nodes or a lower training loss alone do not establish a better player.
Why the last candidate was rejected
The accepted controller scored 75% against Hard on the last fresh bank; the candidate scored 77%. The paired gain was two points, with an estimated 95% interval from −4.5 to +8.5 points. It did not pass the declared threshold. These scores use wins plus half the draws, divided by games.
This was the third hypothesis without an accepted gain, after value-focused training and shared-feature training. Automatic v1 training stopped. The selected checkpoint stayed unchanged; we did not rerun confirmation until a favourable score appeared.
Good enough depends on the product
The models are ready to explain and demonstrate. A perfect Connect Four player is not required for a useful ML Lab. This release therefore provides model notes, measured results and complete Snake recordings.
Live neural browser play has a separate gate: matching encoders and actions, legal moves, fresh browser games, measured response times, cancellation and phone usability. It remains future work. Research should restart with a new hypothesis, a measurable goal and a bounded experiment allowance.