Dungeon With Amnesia
How much does intelligence depend on remembering what is no longer visible?
A PPO laboratory that tests whether agents remember what they can no longer see, by deleting their memory mid-episode.
The setup
At the start of every episode a dying torch reveals a seal, Raven or Moon. The inscription disappears after the first action and can never be observed again. 33 decisions later, at the exit, the agent must invoke the matching seal. The wrong one ends the episode.
Three architectures train with PPO on exactly the same observations: a feed-forward policy with no memory, a policy that sees its last four observations, and a GRU that carries a recurrent state.
Delete the memory, keep everything else
The intervention clones the exact environment and policy at decision 8, then resets the GRU's hidden state in one copy only. Same dungeon, same weights, same action prefix. Below are the real 64-dimensional hidden states from the verified replay (seed 200001). Scrub through the episode.
Decision 0 is the only time the seal is visible. The two agents act identically until the final choice, 33 decisions later. Only the one that kept its memory invokes MOON.
| Intervention | Strength | Control | Treatment | Paired 95% CI of change |
|---|---|---|---|---|
| Reset | at 25 / 50 / 75% | 100% | 43% | −66% to −48% |
| Gaussian noise | σ 0.0 | 100% | 100% | +0% |
| Gaussian noise | σ 0.5 | 100% | 98% | −5% to 0% |
| Gaussian noise | σ 1.0 | 100% | 86% | −21% to −8% |
| Gaussian noise | σ 2.0 | 100% | 67% | −42% to −24% |
Measurements



| Agent | 7 | 11 | 19 | 33 | 53 | 63 decisions |
|---|---|---|---|---|---|---|
| No memory | 57% | 57% | 57% | 57% | 57% | 57% |
| History-4 | 57% | 57% | 57% | 57% | 57% | 57% |
| GRU | 100% | 100% | 100% | 100% | 100% | 100% |
One more section for engineers: architecture, hyper-parameters and design decisions.
Under the hood
Under the hood- Hidden dungeon state
topology, keys, locks, correct seal
- Local observation
58-dim structured encoding; also rendered as text
- Encoder
- MLP: current obs
- History: last 4 obs + actions
- GRU / LSTM: recurrent state
- Policy + value heads
masked discrete actions
- Intervention
- reset
- zero a fraction
- seeded noise
- restore
Recurrent PPO done carefully
Each PPO epoch recomputes the whole episode from zero state under current weights. It never shuffles isolated recurrent timesteps, so backpropagation reaches the original cue. Padding is excluded from losses; time-limit truncations bootstrap the last value.
An explicit information boundary
The encoder never sees coordinates, room IDs, the graph, visited flags, clue truth, the seed or the observer's notebook. A separate observer reconstructs a map for humans only.
Certified solvable
Locks lie on the unique start-to-exit path with keys placed before them. A BFS over (position, inventory, opened locks) certifies every generated dungeon in tests.
Reproducible
Seeded Python/NumPy/PyTorch, deterministic algorithms, one CPU thread, disjoint train/validation/test seed ranges, checkpoint SHA-256 hashes in the report.
What it doesn't do (yet)
- The strong result concerns retaining a one-time cue in a controlled detour, not learned memory of a door's location.
- Quick procedural exploration runs are undertrained and weak; they do not support a generalization claim.
- One optimization seed. Confidence intervals cover map sampling, not training variability.
- Deleting the hidden state also pushes the model off its learned state distribution; it does not prove a localized symbolic memory.