GhostHands

Does imitation transfer a goal-conditioned skill to layouts it has never seen?

Demonstrate a pick-and-place with your hand, change the scene, and test what a learned policy actually generalises.

46% → 97%held-out success of the same MLP when raw poses become relative geometry (100 paired layouts)
Year
2026
Status
Working MVP · simulator only · human data not yet collected
Stack
Python, PyTorch, MediaPipe, JavaScript, FastAPI
01

The idea

Control a virtual gripper with your hand through the webcam: palm position moves it, pushing toward the camera lowers it, turning your hand like a dial rotates it, a pinch closes it. Record complete pick, rotate and place episodes, train a PyTorch policy, then run it in a different layout. During execution you can move the target or the object to test recovery.

The robot side has no scripted fallback. Every action comes from the learned checkpoint and the current scene.

GhostHands interface with live teleoperation on the left and learned execution on the right
02

One attractive rollout is not the evaluation

Task success over 100 layout seeds per suite; the same layouts are used for every policy. CPU run, seed 42, 60 epochs, 240 synthetic training episodes. 1,600 rollouts in total.
PolicyHeld-outWider rangeMoved targetMoved object
Exact replay0%0%0%0%
Absolute BC (MLP)46%36%45%17%
Relative BC (MLP)97%93%98%68%
Recurrent BC (GRU)100%98%99%97%
Bar chart of success with Wilson intervals across four evaluation suites

The interesting jump is representation, not depth: the same MLP goes from 46% to 97% when its inputs become relative geometry instead of raw poses. Mean placement error on held-out layouts drops from 34.7 cm (replay) to 2.1 cm (relative) and 1.2 cm (recurrent), with failed trials included.

One more section for engineers: architecture, hyper-parameters and design decisions.

03

Under the hood

Under the hood
  1. Teleoperation
    • webcam: 21 landmarks + One Euro filters
    • pointer fallback
  2. Deterministic tabletop

    xyz, yaw, binary closure, gravity, attachment constraint

  3. Episode JSON + provenance

    train / validation split by episode

  4. Behaviour cloning
    • MLP 14→128→128→5
    • GRU(96) over 8 frames
  5. Paired evaluation
    • held-out
    • wider range
    • target shift
    • object shift
Observation
14 scalars: relative xy (translation-invariant), heights, yaw as relative sin/cos, grip, contact
Action
5 scalars: motion via tanh + MSE, closure via logit + BCE
Training
AdamW, gradient clipping, deterministic PyTorch, checkpoint chosen on validation loss only
Success
Released open, position error < 5.5 cm, yaw error < 0.25 rad, held for 6 steps
Splits
Train seeds 0–239, validation 10000–10039, each test suite its own seed range
04

What it doesn't do (yet)

  1. The included checkpoints use synthetic teacher demonstrations, not human webcam data. Human recording is implemented but unvalidated with a participant.
  2. A kinematic simulator with exactly known object state: no contact forces, friction, inverse kinematics or perception noise.
  3. A single training seed; the recurrent policy is also larger, so its gains do not isolate temporal context from capacity.

Read the code and the full results on GitHub