Games that teach us about learning.

Two small models. Real experiments. A place to watch what worked, understand what failed, and learn how game-playing AI is built.

Snake

A small neural policy learns expert shortcuts, then receives reinforcement-learning and search refinement. The goal: fill every playable cell, while using fewer movements.

Explore the Snake model

120 / 120 full-board wins

Measured local starts across three 18 × 18 courses. On 60 fresh starts, it used 40.06% fewer movements than the cycle baseline.

Available here: verified recordings with pause, step, speed and seeking. Live neural inference is pending.

Connect Four

A spatial policy/value network learns from self-play and stronger search teachers. Combining the network with a proof controller produces our strongest local player.

Explore the Connect Four model

87.5% against our Expert

The selected network + native search controller, measured on 100 games with varied starts and swapped seats. Draws receive half credit.

Research result. The neural network alone is much weaker. Browser ML play has not been qualified.

The learning journal

Each note follows a question, an experiment and a decision. Failed ideas stay visible because they teach us what to try—and when to stop.

  1. When is the model good enough?

    Three experiments failed to establish a useful gain. We kept the champion and moved effort into a usable learning product.

  2. A stronger teacher, a small unconfirmed gain

    Learning from deeper search improved some scores, but the Hard result was too uncertain to promote.

  3. Lower value loss did not make a stronger player

    A 15.9% improvement in value error did not survive the playing-strength comparison.

  4. Which part of Connect Four is actually learned?

    Separate the neural policy, tree search and exact proofs before interpreting a win rate.

  5. Fill the same board with fewer movements

    The selected Snake policy used 40.06% fewer moves than the cycle baseline on fresh measured starts.

  6. Fix the task before optimizing the model

    Collecting twelve apples was not completing Snake. The real goal is to occupy every playable cell.

A few useful words

Policy
The model’s rule for choosing an action from the current state.
Value
An estimate of how favourable a position or future outcome is.
Self-play
Generating experience by playing games against other versions of an agent.
Warm-start
Initializing a model with useful expert examples before further training.
MCTS / PUCT
Searching possible futures, using policy and value predictions to guide which branches to explore.
Paired evaluation
Comparing two players on the same starting positions to reduce avoidable noise.
Confidence interval
An uncertainty estimate for a measured difference; a positive score difference alone may be inconclusive.
Checkpoint
A saved, fixed set of learned model parameters that can be evaluated reproducibly.