# Selected model robustness: familiar-map specialist with view/recovery weaknesses

The unchanged 85% selected pixel/history model completed 320 fixed-budget diagnostic episodes in 10.59 minutes. No optimizer, teacher, restore, extra life or rescue ran. Model inputs remained RGB, its actual previous command and learned memory. Pose and terminal data were logs only.

| Condition | Exits | Deaths | Timeouts | 95% Wilson completion interval |
|---|---:|---:|---:|---:|
| Ordinary MAP01 | **84/100** | 12 | 4 | 75.6–89.9% |
| Same seeds, RGB 25% darker | **60/100** | 25 | 15 | 50.2–69.1% |
| Changed starting heading | **29/40** | 6 | 5 | 57.2–83.9% |
| Movement disturbance, then neural recovery | **19/40** | 12 | 9 | 32.9–62.5% |
| Unfamiliar MAP02 | **0/40** | 38 | 2 | 0–8.8% |

Darker images lost 29 normally successful seeds and gained five, a net loss of 24 exits. This intervention changed only actor-visible RGB, not engine rules. Normal 84% is consistent with the independently confirmed 85% baseline; neither result meets the original 90% goal.

The heading and movement conditions share the first 40 normal seeds, on which ordinary play scored 31/40. Eight initial left/right commands changed heading by median 47.46 degrees; 29/40 thereafter exited, with seven perturbation-only and nine normal-only exits. The movement condition played 128 neural commands, then 32 seeded random left/right/forward commands, then returned to neural control. All 40 survived the disturbance and moved at least 32 map units; median displacement was 157.41. It lost 14 normally successful seeds and gained two. Learned memory processed every image and received actual executed commands during disturbances. All setup commands consumed the original budget. These are explicitly perturbed episodes, not autonomous ordinary-start qualification; no setup failures were excluded.

The MAP02 result limits the model to familiar-map performance in this evidence. One different level changes layout, enemies and encounters together; it does not isolate a cause or prove that transfer is impossible. Fresh MAP01 RNG seeds are not unfamiliar geometry.

Independent checks verified every trace hash, intervention schedule, actual/suggested command, denominator, budget, summary and confidence interval. Fourteen native neural replays covered every outcome present in each condition and reproduced raw/actor-visible RGB, proposals, executed commands, poses, timing and terminal metrics. The separate five-condition smoke also passed five native replays. All 196 Python tests passed. The selected checkpoint and source/IWAD hashes remained unchanged; final campaign qualification seeds remain untouched.

Next prioritize visual invariance and recovery training on fresh fitting-only trajectories, with normal-view behavior preserved and a matched control. Register any new hypothesis before fitting; do not retune the closed combat-gate recipe. Darker-view failure suggests sensitivity to pixels; movement failure suggests a recovery weakness. Neither uniquely proves a feature or memory capacity limit. Independent lineages and ordinary-start campaign qualification remain required. No job is running.
