Skip to content
← All experiments

Learning a level. Learning when to stop.

What two days of visual-control research taught us about demonstrations, memory, evaluation and diminishing returns. Watch the selected model complete MAP01 and see where it fails.

Freedoom · Research recorded 8–10 October 2026

Actual selected-model gameplay · successful MAP01 example

1. Make completion the goal

After Snake and Connect Four, we tried visual control in a 3D game. Freedoom supplies freely licensed game content; ViZDoom lets a model interact with the native engine. Our target was to finish Freedoom 2 MAP01 at skill 2 using a screen image and learned memory, without a route planner choosing moves during play.

A real exit counts as success. Deaths and timeouts count as failures, with one ordinary start and the original 3,500-tic budget. The target was at least 90% normal-start completion, independent confirmation and training replication, with separate robustness reports.

2. Watch an actual successful run

66 seconds, silent. Full real-time gameplay followed by a short exit receipt. Seed 846400: the first successful game in the confirmation bank, not a search for the fastest run.

The model exits after 1,080 two-tic commands, about 61.7 seconds, with 94 health remaining. A fresh native-engine replay matched all 1,081 recorded screen states. The observation stream and terminal outcome also matched the previously audited confirmation run.

This is one successful example. It does not show every failure or imply that the model always wins. Download the video (21.5 MiB). Game visuals are from Freedoom; content credits and licence.

3. See what the network actually receives

The actor receives one 160×120 RGB image, its previous executed action and a learned 128-unit recurrent memory. A recurrent network carries information between decisions: the current image may not reveal which doorway it just passed or whether a turn is still useful.

The six choices are idle, turn left, turn right, move forward, fire and use. Each command runs for two native game tics. The selected policy samples its learned action probabilities at temperature 1. Temperature changes how concentrated sampling is; it does not add planning.

The model combines neural navigation and combat predictions through a learned mixing gate. Coordinates, enemy annotations and engine state can support training targets or audit logs, but do not enter the actor at evaluation. There is no runtime teacher, map solver, state restoration or extra life.

4. Use demonstrations, then test improvements

  1. Learn from examplesTeach useful navigation decisions
  2. Preserve combat behaviorAvoid breaking an existing skill
  3. Freeze and compareJudge real exits on fresh games
Training assistance creates examples; the frozen player must act on its own during evaluation.

Imitation learning teaches a model to match examples from a stronger teacher. Reinforcement learning changes behavior using outcomes from interaction. We explored both, but the selected campaign improvement is supervised imitation and preservation; it is not a new RL success.

The selected continuation fitted a 64-unit navigation decoder with 4,000 supervised updates. A separate preservation loss kept the earlier reference’s combat-related behavior from being overwritten. The corpus contained 530 games, split into 424 fitting games and 106 held-out games. Normalization used fitting records only; held-out and evaluation games did not enter fitting.

The navigation change alone was weaker than the preserved version. On the independent confirmation, the previous reference exited 281/400 games, the navigation-only control 300/400, and the preserved candidate 340/400. A control is a deliberately matched comparison that helps explain which part of a change mattered.

Separate easier exercises did achieve genuine RL results: the move-and-shoot task and a demonstration-initialized maze policy. The selected maze policy passed 96/100 starts, while a second independent training lineage reached 88/100 and missed its 90% gate. Those results do not qualify the full campaign or prove the training recipe is reliable.

5. Measure strength and its limits

Frozen selected model · native local evaluation · real exits / all starts
ConditionExitsWhat it tests
Independent MAP01 confirmation340 / 400 (85%)Familiar-level completion
Separate ordinary MAP01 diagnostic84 / 100Another fresh random-seed bank
Same diagnostic seeds, 25% darker RGB60 / 100View sensitivity
Changed starting heading29 / 40Heading disturbance
Movement disturbance, then neural control19 / 40Recovery after leaving its route
Unfamiliar MAP020 / 40Transfer to another level

The 400-game confirmation includes 45 deaths and 15 timeouts. Its 95% Wilson completion interval is 81.2–88.2%. An interval expresses sampling uncertainty on this evaluation bank; it is not a guarantee for another map or a new training run.

The heading and movement tests use the first 40 ordinary diagnostic seeds, where normal play exited 31/40. Disturbance commands consumed the original budget, and memory received the actual executed actions. These are recovery tests, not ordinary-start qualification.

6. Keep the experiments that did not help

Later candidates · matched normal-view comparisons within each trial
AttemptCandidate / reference exitsDecision
Spatial-feature fitting164/200 vs 167/200Rejected; matched control 169/200
Combat-head terminal-reward RL162/200 vs 163/200Rejected
Learned RGB encoder32/40 vs 35/40Rejected at pilot gate
Full-actor RL33/40 vs 34/40Rejected at pilot gate

Each row is a separate paired experiment with its own seed bank. Compare candidate with reference inside a row; do not rank percentages across different banks. Small differences alone do not prove superiority. Failed pilots did not proceed to larger confirmation or replace the selected checkpoint.

A learned value baseline also missed its predeclared variance-reduction gate. Fixing gradients made full-actor training update the intended modules, but the resulting player still failed its promotion test. An implementation repair and a stronger model are different outcomes.

Counterfactual work asked what would happen if one action changed and the model then resumed normal play. Reused-engine and saved-state attempts failed reproducibility checks; fresh engines with exact prefix replay produced repeatable forks. That established a training-data mechanism, not a stronger actor. The proposed utility screen was not run.

7. Carry these lessons into the next model

Reward is a proxy; the goal is an outcome.
Collecting rewards, fitting teacher actions or reporting changing-policy training wins does not establish that a frozen player completes the level more reliably. Evaluate real exits.
Preserve old skills while teaching new ones.
Improving navigation can damage combat. The preservation comparison showed that matching an existing skill can matter as much as fitting the new one.
Investigate failures before adding capacity.
Darker images and route disturbances reveal weaknesses. They suggest targeted hypotheses, but do not uniquely diagnose network size, perception or memory as the cause.
Verify the simulator and the full player.
Replay hashes, actual action history, terminal checks and original time budgets caught issues that a loss curve could not. A model can learn from a broken measurement pipeline.
Use papers as hypotheses, not promises.
A published method motivates a controlled trial. Adapting one idea on our data and local budget does not reproduce the paper’s benchmark or guarantee its gains.
Keep evaluation sealed and failures visible.
Freeze candidates before fresh comparisons. Avoid tuning on a failed confirmation bank. Multiple continuations from one parent do not establish independent training reliability.

The same pattern appeared in Connect Four: lower value loss did not establish better playing strength. Snake taught us to fix the actual full-board objective before optimizing movements. Across all three, define what success means, then make the measurements match it.

8. Stop when the next run stops being a good experiment

After more than two days, several distinct later candidates failed their promotion gates. We retained the 85% checkpoint and paused open-ended research. The engineering became more trustworthy, but the recent work did not produce an accepted playing-strength gain.

Good enough for a learning article: a genuine playthrough, reproducible measurements, clear method labels and visible limitations. Good enough for the original reliable-player target: not yet. We keep the 90% threshold, independent confirmation and training-lineage requirements rather than lowering them after seeing the scores.

Future research should start with one specific hypothesis, one preregistered candidate and at most four hours of local compute including evaluation. A paired 200-game gate precedes independent confirmation only if it passes. Stop when the gate fails or the budget ends; do not keep changing the recipe on that same bank. An exhausted budget means the research is incomplete, not successful.

9. Preserve evidence and make the next step useful

The selected confirmation audited 1,200 action traces across the candidate and two comparisons, plus nine native neural replays. The final campaign seeds 900000–900099 remain untouched. The campaign selection is one training lineage, not three independent full training replications.

Model identity and reproduction

Selected direct-navigation expert with combat preservation. Checkpoint SHA-256:

61e1f7d8b08f999ce8439c0f04036c56888afac9777c5c95309baef02fb65cbc

Local training and audit scripts live under experiments/rl/freedoom. Reproduction requires pinned ViZDoom assets and local checkpoints/data; those are not all distributed with this article. The video is a verified recording, not browser inference.

Source reports

Research behind the exercises

ViZDoom provides a visual RL platform. Proximal Policy Optimization informed our separate recurrent RL exercises. Sample Factory highlights training throughput; its published compute setup is not our local benchmark.

The next useful product step is this read/watch account. Live browser gameplay would require a separate engine/runtime integration, inference parity and mobile performance qualification.