# Playthrough and stopping recommendation — 10 October 2026

The selected MAP01 specialist is a useful educational demo. It is not qualified against the original full goal, and recent trials have not improved the selected 340/400 (85%) result. Recommend ending open-ended training now, preserving the checkpoint and evidence. The user has requested a stopping assessment; the ongoing goal is now paused.

## Delivered playthrough

Site video: `/ml/videos/freedoom-map01.mp4`.
Final video: `/ml/videos/freedoom-map01.mp4` (66 seconds, 1280×720, 30 fps). Raw gameplay: `native/gameplay.mp4`. Seed 846400, the first successful independent confirmation row. 1080 commands, each two native tics, about 61.714 seconds. Real MAP01 skill 2 level exit, health 94, seven kills. No cuts, speed change, runtime teacher, route planner, state restoration or extra lives. Silent capture. RGB + previous executed own action + learned recurrent memory, categorical temperature 1, RNG seed+3129. Selected checkpoint SHA256 `61e1f7d8b08f999ce8439c0f04036c56888afac9777c5c95309baef02fb65cbc`.

`record_selected.py` independently regenerated the exact previously audited observation stream and terminal telemetry. All 1081 screen states matched a fresh recorded-engine replay, which also reached a genuine exit. `freedoom-playthrough-receipt.json` preserves footage/demo/assets/checkpoint hashes. A successful example is explicitly labelled; it does not imply 100% completion.

## Why stop this run

- Frozen selected model: 340/400 exits, 45 deaths, 15 timeouts. Latest candidates remain unpromoted.
- Frozen-feature spatial fitting: 164/200 versus 167 reference and 169 control; rejected.
- Combat-head terminal-reward RL: 162/200 versus 163 reference; rejected after about 56.5 minutes.
- Learned RGB encoder: 32/40 normal and 12/20 darker versus reference 35/40 and 15/20; rejected.
- Full-actor RL after correcting training gradients: 33/40 normal and 13/20 darker versus reference 34/40 and 13/20; rejected after about 60.9 minutes plus verification.
- Critic-baseline screening did not meet its preregistered variance reduction gate.
- Counterfactual work established reproducible simulator forks, not a stronger actor. Utility screen code exists but has not been registered, tested or run.

The engineering and verification work made the results more trustworthy. It has not established further playing-strength gains. More iterations are possible, but there is no evidence that indefinite local retries will reach the target efficiently.

## What is good enough

For a learning article and a clearly labelled local MAP01 demo: good enough now. Method is supervised imitation/preservation; do not describe selected weights as an RL breakthrough. On unfamiliar MAP02 it exits 0/40; darker and off-route performance also degrades.

For the original reliable-player goal: not good enough. Keep the >=90% normal-exit threshold, independent confirmation, robustness/layout reports and independent training lineage requirements. Do not lower them retroactively. Final seeds 900000–900099 remain untouched. No browser integration or release is part of this video task.

If research resumes, treat it as a new bounded experiment, not an endless continuation: one specific hypothesis, one preregistered candidate, at most four hours of local compute including evaluation, a paired 200-game gate followed by independent confirmation only if it passes. Stop on a failed gate or budget exhaustion; no post-result tuning on that bank. Budget exhaustion means incomplete, not achieved.

Final MP4 verified: 66.000 seconds, 1980 H.264 frames at 30 fps, 1280×720, 22,553,198 bytes. HyperFrames check passed with zero findings and 61/61 contrast checks. Early/middle/late/exit frames were inspected from the encoded MP4. Final artifact hash is in the video project delivery receipt.

Freedoom game-content notice: ../videos/freedoom-COPYING.adoc (upstream COPYING.adoc, retrieved 10 October 2026). No game engine or IWAD is distributed in this article; the assets are recorded gameplay and screenshots.
