Dungeon With Amnesia

How much does intelligence depend on remembering what is no longer visible?

A PPO laboratory that tests whether agents remember what they can no longer see, by deleting their memory mid-episode.

100 → 43escapes out of 100 when the GRU's hidden state is reset mid-episode
Year
2026
Status
Research demo · one training seed · results reported with intervals
Stack
Python, PyTorch, Streamlit, NumPy
01

The setup

At the start of every episode a dying torch reveals a seal, Raven or Moon. The inscription disappears after the first action and can never be observed again. 33 decisions later, at the exit, the agent must invoke the matching seal. The wrong one ends the episode.

Three architectures train with PPO on exactly the same observations: a feed-forward policy with no memory, a policy that sees its last four observations, and a GRU that carries a recurrent state.

100/100GRU escapes on held-out showcase maps95% Wilson CI 96.3–100%
57/100no-memory and history-4 agents: chance on a binary choice
30/30counterfactual pairs where flipping only the cue flipped the GRU's final choice
02

Delete the memory, keep everything else

The intervention clones the exact environment and policy at decision 8, then resets the GRU's hidden state in one copy only. Same dungeon, same weights, same action prefix. Below are the real 64-dimensional hidden states from the verified replay (seed 200001). Scrub through the episode.

Verified replay · seed … · memory reset after decision …64 hidden units × … decisions
Control · memory intact
action —Failed
Treatment · hidden state reset
action —Wrong seal

Decision 0 is the only time the seal is visible. The two agents act identically until the final choice, 33 decisions later. Only the one that kept its memory invokes MOON.

00/0
Paired interventions over 100 test maps. Positive control: zero noise must preserve behaviour.
InterventionStrengthControlTreatmentPaired 95% CI of change
Resetat 25 / 50 / 75%100%43%−66% to −48%
Gaussian noiseσ 0.0100%100%+0%
Gaussian noiseσ 0.5100%98%−5% to 0%
Gaussian noiseσ 1.0100%86%−21% to −8%
Gaussian noiseσ 2.0100%67%−42% to −24%
03

Measurements

PPO training curves for the three architectures
Cue-to-decision distance sweep. Every delay exceeds the history-4 window, so this sweep does not locate that window's boundary.
Agent71119335363 decisions
No memory57%57%57%57%57%57%
History-457%57%57%57%57%57%
GRU100%100%100%100%100%100%

One more section for engineers: architecture, hyper-parameters and design decisions.

04

Under the hood

Under the hood
  1. Hidden dungeon state

    topology, keys, locks, correct seal

  2. Local observation

    58-dim structured encoding; also rendered as text

  3. Encoder
    • MLP: current obs
    • History: last 4 obs + actions
    • GRU / LSTM: recurrent state
  4. Policy + value heads

    masked discrete actions

  5. Intervention
    • reset
    • zero a fraction
    • seeded noise
    • restore

Recurrent PPO done carefully

Each PPO epoch recomputes the whole episode from zero state under current weights. It never shuffles isolated recurrent timesteps, so backpropagation reaches the original cue. Padding is excluded from losses; time-limit truncations bootstrap the last value.

An explicit information boundary

The encoder never sees coordinates, room IDs, the graph, visited flags, clue truth, the seed or the observer's notebook. A separate observer reconstructs a map for humans only.

Certified solvable

Locks lie on the unique start-to-exit path with keys placed before them. A BFS over (position, inventory, opened locks) certifies every generated dungeon in tests.

Reproducible

Seeded Python/NumPy/PyTorch, deterministic algorithms, one CPU thread, disjoint train/validation/test seed ranges, checkpoint SHA-256 hashes in the report.

05

What it doesn't do (yet)

  1. The strong result concerns retaining a one-time cue in a controlled detour, not learned memory of a door's location.
  2. Quick procedural exploration runs are undertrained and weak; they do not support a generalization claim.
  3. One optimization seed. Confidence intervals cover map sampling, not training variability.
  4. Deleting the hidden state also pushes the model off its learned state distribution; it does not prove a localized symbolic memory.

Read the code and the full results on GitHub