GhostHands
Does imitation transfer a goal-conditioned skill to layouts it has never seen?
Demonstrate a pick-and-place with your hand, change the scene, and test what a learned policy actually generalises.
The idea
Control a virtual gripper with your hand through the webcam: palm position moves it, pushing toward the camera lowers it, turning your hand like a dial rotates it, a pinch closes it. Record complete pick, rotate and place episodes, train a PyTorch policy, then run it in a different layout. During execution you can move the target or the object to test recovery.
The robot side has no scripted fallback. Every action comes from the learned checkpoint and the current scene.

One attractive rollout is not the evaluation
| Policy | Held-out | Wider range | Moved target | Moved object |
|---|---|---|---|---|
| Exact replay | 0% | 0% | 0% | 0% |
| Absolute BC (MLP) | 46% | 36% | 45% | 17% |
| Relative BC (MLP) | 97% | 93% | 98% | 68% |
| Recurrent BC (GRU) | 100% | 98% | 99% | 97% |

The interesting jump is representation, not depth: the same MLP goes from 46% to 97% when its inputs become relative geometry instead of raw poses. Mean placement error on held-out layouts drops from 34.7 cm (replay) to 2.1 cm (relative) and 1.2 cm (recurrent), with failed trials included.
One more section for engineers: architecture, hyper-parameters and design decisions.
Under the hood
Under the hood- Teleoperation
- webcam: 21 landmarks + One Euro filters
- pointer fallback
- Deterministic tabletop
xyz, yaw, binary closure, gravity, attachment constraint
- Episode JSON + provenance
train / validation split by episode
- Behaviour cloning
- MLP 14→128→128→5
- GRU(96) over 8 frames
- Paired evaluation
- held-out
- wider range
- target shift
- object shift
- Observation
- 14 scalars: relative xy (translation-invariant), heights, yaw as relative sin/cos, grip, contact
- Action
- 5 scalars: motion via tanh + MSE, closure via logit + BCE
- Training
- AdamW, gradient clipping, deterministic PyTorch, checkpoint chosen on validation loss only
- Success
- Released open, position error < 5.5 cm, yaw error < 0.25 rad, held for 6 steps
- Splits
- Train seeds 0–239, validation 10000–10039, each test suite its own seed range
What it doesn't do (yet)
- The included checkpoints use synthetic teacher demonstrations, not human webcam data. Human recording is implemented but unvalidated with a participant.
- A kinematic simulator with exactly known object state: no contact forces, friction, inverse kinematics or perception noise.
- A single training seed; the recurrent policy is also larger, so its gains do not isolate temporal context from capacity.