Reward Goblin

If an agent only ever sees a number, what does it learn when the number is slightly wrong?

A physics playground where real PPO agents optimise the reward you wrote instead of the task you meant.

53.5 vs 3.9reward earned by the cheating policy vs the correct solution (Edge Goblin v1, seed 1)
Year
2026
Status
Working research demo · 7 experiments · 48 trained runs
Stack
Python, PyTorch, Stable-Baselines3, Gymnasium, Pymunk, FastAPI, JavaScript
01

The problem

A reinforcement-learning agent never sees your intention, only a scalar reward. If the reward can be maximised without doing the task, a competent optimiser will find that way. The literature calls it specification gaming. Most people only ever read about it.

Reward Goblin makes it measurable. You describe “push the box into the exit” as a reward function, a real PPO policy trains on it, and the interface replays what it learned next to a scripted solution on the identical start state, with the reward each one earned.

The agent does not do what you intended. It does what you rewarded.

02

Watch it happen

The replays below are drawn live from the recorded episode files committed to the repository: positions, rewards and detector events, frame by frame. No GIF, no scripting. Left is the hand-written reference controller; right is the trained policy.

loading replay…
What you intendedscripted A*
What you rewardedPPO policy

000/0
03

Seven goblins, seven loopholes

Each row is a real PPO run (3 seeds × 800k steps), evaluated deterministically on held-out start states. Nothing below was scripted.
GoblinThe reward (v1)What PPO actually learned
Edge+0.2 every step the box touches the exit; episode ends once fully insidePushes the box until it overlaps the exit, then stops. Fully inside would end the income.
Touch+5 each time the box starts touching the exitOne seed of three learned to push the box in and out repeatedly (95% of its episodes).
Distance+1 per metre closer; retreating is freeJiggles the box: one replay was paid for 7.1 m of progress while the box ended 2.4 m closer.
SpeedReward for moving toward the exit quicklyIgnores the box and runs laps through the exit.
Survival+0.1 per step alive, +5 for completionParks the box next to the exit and waits out the clock.
Lava−0.05 per step; lava −1 and ends the episodeWalks straight into the lava: −1 is cheaper than 300 steps of time penalty.
Wall+0.1 per step while the box is within 2.5 m (straight line)Pins the box against the outside of the exit room's wall.
04

Measuring the real task separately

The human objective lives in its own module, true_objective.py, and no training reward ever reads it. When an episode ends, for whatever reason, the goblin is frozen and physics keeps running for 30 more frames; the box must stay fully inside the exit the whole time. That separates “resting in the exit” from “sliding through it when the clock ran out”.

Every reward version trains on 3 seeds and is evaluated on 20 held-out start states and 20 procedurally generated layouts the policy never saw. One clean replay proves nothing.

0–5%true success of every misspecified v1 reward
80%true success after rewarding the outcome (v2), 0% exploits
28%the same v2 policies on unseen layouts
~300ksteps for the edge exploit to saturate; the honest reward is still improving at 800k
Bar chart of true task success versus exploit rate for every reward version
Training curves: exploit frequency saturates early for v1 while v2 true success keeps rising
05

What the data says

The optimiser isn't failing; the objective is.

In all 7 experiments the cheating policy's mean return beats the scripted solution under the same reward: 20.1 vs 4.1 (Edge), 43.5 vs 16.3 (Speed), 24.9 vs 15.8 (Wall).

A fix on the training map is not a fix.

Policies that reach 80% on their training map succeed on 28% of unseen layouts, while the scripted reference solves all of them.

Seeds matter.

Touch v1's farming loop was found by 1 seed of 3. Wall v3 solved 95% of episodes for seed 3 and 0% for seeds 1–2. A single run would have told either story.

Removing an exploit is not teaching the task.

Dropping the proximity bonus removed wall-pinning but left 0% success, because straight-line shaping still points into the wall.

One more section for engineers: architecture, hyper-parameters and design decisions.

06

Under the hood

Under the hood
  1. Reward editor
    • JSON components
    • static linter

    No code path executes user input

  2. FastAPI
    • REST API
    • job manager · process pool
  3. Training
    • SB3 PPO / A2C
    • metrics callback
  4. Environment
    • Gymnasium env
    • Pymunk physics
    • reward components
  5. Analysis
    • true objective
    • exploit detectors
    • A* reference
  6. Replay UI
    • split-screen canvas
    • timeline + markers

Everything the UI shows is read from files that training writes: run metadata, per-rollout metrics, evaluations and episode recordings.

Environment
Gymnasium API, passes SB3 check_env
Physics
Pymunk; velocity-servoed goblin, box with floor friction, rotation locked
Control
15 Hz, 4 physics substeps per action, 300-step episodes
Actions
Discrete(5): no-op, left, right, up, down
Observations
37 floats: positions, velocities, relative vectors, contact flags, time remaining, 8 wall rays, 8 lava rays
Algorithm
PPO, MLP 64×64, 8 envs, n_steps 512, batch 256, 10 epochs, γ 0.99, λ 0.95, entropy 0.01, clip 0.2
Budget
800k steps × 3 seeds per reward version; 48 runs ≈ 2.5 h on 10 CPU cores
Tests
55: env checker, rewards, detectors, API, replays

Before any training, a static reward linter warns which loopholes a reward is likely to open, for example that touching lava may be cheaper than surviving a full episode under the step penalty. Each run stores algorithm, hyper-parameters, seed, environment version, reward fingerprint, library versions and git commit, and same seeds reproduce bit-identical results.

07

What it doesn't do (yet)

  1. Exploit detectors are hand-tuned heuristics: diagnostics, not proofs. Expect false negatives on rewards they were not tuned for.
  2. Policies that learn the task on their training map succeed on only 28% of unseen procedural layouts.
  3. The scripted A* reference is not perfect either: 37–40 of 40 per layout type in a spot check. Its failures are reported.

Read the code and the full results on GitHub