A small neural policy learns expert shortcuts, then receives reinforcement-learning and search refinement. The goal: fill every playable cell, while using fewer movements.
Measured local starts across three 18 × 18 courses. On 60 fresh starts, it used 40.06% fewer movements than the cycle baseline.
Available here: verified recordings with pause, step, speed and seeking. Live neural inference is pending.
Connect Four
A spatial policy/value network learns from self-play and stronger search teachers. Combining the network with a proof controller produces our strongest local player.