1. Make completion the goal
After Snake and Connect Four, we tried visual control in a 3D game. Freedoom supplies freely licensed game content; ViZDoom lets a model interact with the native engine. Our target was to finish Freedoom 2 MAP01 at skill 2 using a screen image and learned memory, without a route planner choosing moves during play.
A real exit counts as success. Deaths and timeouts count as failures, with one ordinary start and the original 3,500-tic budget. The target was at least 90% normal-start completion, independent confirmation and training replication, with separate robustness reports.
2. Watch an actual successful run
The model exits after 1,080 two-tic commands, about 61.7 seconds, with 94 health remaining. A fresh native-engine replay matched all 1,081 recorded screen states. The observation stream and terminal outcome also matched the previously audited confirmation run.
This is one successful example. It does not show every failure or imply that the model always wins. Download the video (21.5 MiB). Game visuals are from Freedoom; content credits and licence.
3. See what the network actually receives
The actor receives one 160×120 RGB image, its previous executed action and a learned 128-unit recurrent memory. A recurrent network carries information between decisions: the current image may not reveal which doorway it just passed or whether a turn is still useful.
The six choices are idle, turn left, turn right, move forward, fire and use. Each command runs for two native game tics. The selected policy samples its learned action probabilities at temperature 1. Temperature changes how concentrated sampling is; it does not add planning.
The model combines neural navigation and combat predictions through a learned mixing gate. Coordinates, enemy annotations and engine state can support training targets or audit logs, but do not enter the actor at evaluation. There is no runtime teacher, map solver, state restoration or extra life.
4. Use demonstrations, then test improvements
- Learn from examplesTeach useful navigation decisions
- Preserve combat behaviorAvoid breaking an existing skill
- Freeze and compareJudge real exits on fresh games
Imitation learning teaches a model to match examples from a stronger teacher. Reinforcement learning changes behavior using outcomes from interaction. We explored both, but the selected campaign improvement is supervised imitation and preservation; it is not a new RL success.
The selected continuation fitted a 64-unit navigation decoder with 4,000 supervised updates. A separate preservation loss kept the earlier reference’s combat-related behavior from being overwritten. The corpus contained 530 games, split into 424 fitting games and 106 held-out games. Normalization used fitting records only; held-out and evaluation games did not enter fitting.
The navigation change alone was weaker than the preserved version. On the independent confirmation, the previous reference exited 281/400 games, the navigation-only control 300/400, and the preserved candidate 340/400. A control is a deliberately matched comparison that helps explain which part of a change mattered.
Separate easier exercises did achieve genuine RL results: the move-and-shoot task and a demonstration-initialized maze policy. The selected maze policy passed 96/100 starts, while a second independent training lineage reached 88/100 and missed its 90% gate. Those results do not qualify the full campaign or prove the training recipe is reliable.
5. Measure strength and its limits
| Condition | Exits | What it tests |
|---|---|---|
| Independent MAP01 confirmation | 340 / 400 (85%) | Familiar-level completion |
| Separate ordinary MAP01 diagnostic | 84 / 100 | Another fresh random-seed bank |
| Same diagnostic seeds, 25% darker RGB | 60 / 100 | View sensitivity |
| Changed starting heading | 29 / 40 | Heading disturbance |
| Movement disturbance, then neural control | 19 / 40 | Recovery after leaving its route |
| Unfamiliar MAP02 | 0 / 40 | Transfer to another level |
The 400-game confirmation includes 45 deaths and 15 timeouts. Its 95% Wilson completion interval is 81.2–88.2%. An interval expresses sampling uncertainty on this evaluation bank; it is not a guarantee for another map or a new training run.
The heading and movement tests use the first 40 ordinary diagnostic seeds, where normal play exited 31/40. Disturbance commands consumed the original budget, and memory received the actual executed actions. These are recovery tests, not ordinary-start qualification.
6. Keep the experiments that did not help
| Attempt | Candidate / reference exits | Decision |
|---|---|---|
| Spatial-feature fitting | 164/200 vs 167/200 | Rejected; matched control 169/200 |
| Combat-head terminal-reward RL | 162/200 vs 163/200 | Rejected |
| Learned RGB encoder | 32/40 vs 35/40 | Rejected at pilot gate |
| Full-actor RL | 33/40 vs 34/40 | Rejected at pilot gate |
Each row is a separate paired experiment with its own seed bank. Compare candidate with reference inside a row; do not rank percentages across different banks. Small differences alone do not prove superiority. Failed pilots did not proceed to larger confirmation or replace the selected checkpoint.
A learned value baseline also missed its predeclared variance-reduction gate. Fixing gradients made full-actor training update the intended modules, but the resulting player still failed its promotion test. An implementation repair and a stronger model are different outcomes.
Counterfactual work asked what would happen if one action changed and the model then resumed normal play. Reused-engine and saved-state attempts failed reproducibility checks; fresh engines with exact prefix replay produced repeatable forks. That established a training-data mechanism, not a stronger actor. The proposed utility screen was not run.
7. Carry these lessons into the next model
- Reward is a proxy; the goal is an outcome.
- Collecting rewards, fitting teacher actions or reporting changing-policy training wins does not establish that a frozen player completes the level more reliably. Evaluate real exits.
- Preserve old skills while teaching new ones.
- Improving navigation can damage combat. The preservation comparison showed that matching an existing skill can matter as much as fitting the new one.
- Investigate failures before adding capacity.
- Darker images and route disturbances reveal weaknesses. They suggest targeted hypotheses, but do not uniquely diagnose network size, perception or memory as the cause.
- Verify the simulator and the full player.
- Replay hashes, actual action history, terminal checks and original time budgets caught issues that a loss curve could not. A model can learn from a broken measurement pipeline.
- Use papers as hypotheses, not promises.
- A published method motivates a controlled trial. Adapting one idea on our data and local budget does not reproduce the paper’s benchmark or guarantee its gains.
- Keep evaluation sealed and failures visible.
- Freeze candidates before fresh comparisons. Avoid tuning on a failed confirmation bank. Multiple continuations from one parent do not establish independent training reliability.
The same pattern appeared in Connect Four: lower value loss did not establish better playing strength. Snake taught us to fix the actual full-board objective before optimizing movements. Across all three, define what success means, then make the measurements match it.
8. Stop when the next run stops being a good experiment
After more than two days, several distinct later candidates failed their promotion gates. We retained the 85% checkpoint and paused open-ended research. The engineering became more trustworthy, but the recent work did not produce an accepted playing-strength gain.
Good enough for a learning article: a genuine playthrough, reproducible measurements, clear method labels and visible limitations. Good enough for the original reliable-player target: not yet. We keep the 90% threshold, independent confirmation and training-lineage requirements rather than lowering them after seeing the scores.
Future research should start with one specific hypothesis, one preregistered candidate and at most four hours of local compute including evaluation. A paired 200-game gate precedes independent confirmation only if it passes. Stop when the gate fails or the budget ends; do not keep changing the recipe on that same bank. An exhausted budget means the research is incomplete, not successful.
9. Preserve evidence and make the next step useful
The selected confirmation audited 1,200 action traces across the candidate and two comparisons, plus nine native neural replays. The final campaign seeds 900000–900099 remain untouched. The campaign selection is one training lineage, not three independent full training replications.
Model identity and reproduction
Selected direct-navigation expert with combat preservation. Checkpoint SHA-256:
61e1f7d8b08f999ce8439c0f04036c56888afac9777c5c95309baef02fb65cbcLocal training and audit scripts live under experiments/rl/freedoom. Reproduction requires pinned ViZDoom assets and local checkpoints/data; those are not all distributed with this article. The video is a verified recording, not browser inference.
Source reports
- Selected results, experiment decisions and stopping policy
- Robustness protocol, results and interpretation
- Playthrough verification and model identity
Research behind the exercises
ViZDoom provides a visual RL platform. Proximal Policy Optimization informed our separate recurrent RL exercises. Sample Factory highlights training throughput; its published compute setup is not our local benchmark.
The next useful product step is this read/watch account. Live browser gameplay would require a separate engine/runtime integration, inference parity and mobile performance qualification.
