Which part of Connect Four is actually learned?

Separate the neural policy, tree search and exact proofs before interpreting a win rate.

· Leaf Arcade experiment notes

Three players from one network

The raw network ranks the seven columns and estimates the position’s value. Legal-move masking prevents choosing a full column. Its engineered inputs also encode immediate wins and opponent threats.

Adding 256 PUCT simulations makes a different player: the network guides a tree of possible continuations. The accepted local controller adds bounded native search and exact proofs as well. Each combination needs its own benchmark.

What the measurements say

On the recorded full-controller bank, the selected system scored 87.5% against our custom Expert and 70% against Fairy Hard. The larger MCTS-only bank scored 76.5% against Expert and 36% against Hard. Raw neural play scored 29.2% and 19.2% respectively on a smaller bank.

Those are different evaluation banks and sample sizes, not a controlled paired comparison between modes. They show why serving only the neural file must not inherit the full controller’s headline score. In one controller audit, exact proofs handled 93.1% of moves.

How the model learned

This iteration uses search-supervised policy and value learning on positions from self-play and opponents. Expert search supplies teaching targets. It follows the useful pattern of learning a policy/value model to support planning, rather than pretending every strategic rule emerged from scratch.

The accepted network contains 69,102 learned parameters. Its checkpoint and ONNX export are frozen. A matching browser implementation and its own qualification are still pending.

Ideas to read next

These sources informed the method. Our small experiments do not reproduce their full research results.