Telekinesis
Can you pick up anything on your desk, in video, without a class list or a 3D model?
Touch a real object in the webcam image, pinch, and lift its appearance out of the video while the background is reconstructed behind it.
One interaction, eight frames
These frames come from running the project's own fixture script: real EdgeSAM, real app logic, a drawn hand driving synthetic landmarks over a synthetic desk. Nothing is mocked except the camera.

Real pipeline output on the repo's synthetic desk fixture. No webcam footage.
Detection, segmentation and tracking are different jobs
Detection
MediaPipe finds hands. No object detector and no class list: anything coherent can be picked up.
Segmentation
EdgeSAM, a promptable SAM-family model, answers “what object is at this point?” once per selection, in a background thread. The prompt sits slightly ahead of the fingertip, because the fingertip pixel is skin; points on the hand go in as negative prompts.
Choosing the whole object
SAM's own score favours small crisp parts. Ranking drops implausible masks, then prefers the largest candidate whose score and stability (does the mask change if the logit threshold moves ±1?) are close to the best. Stability is what rejects merges of neighbouring objects.
Tracking
Masked normalised cross-correlation in a small window follows the same object every frame. It can report OCCLUDED or LOST and hold still; it never switches objects.
One more section for engineers: architecture, hyper-parameters and design decisions.
Under the hood
Under the hood- Transforms
- BGR→RGB, SAM normalisation, longest side 1024, pad; logits upsampled, padding cropped, then resized back
- Compositing
- Feathered soft alpha; premultiplied colour before warping to avoid dark fringes
- Geometry
- One affine T(pos)·R(angle)·S(scale)·T(−anchor) applied to pixels and alpha alike
- Reconstruction
- Clean plate → scene memory (last ~40 s) → pyramid fill, with OpenCV FSR in the background
- Occlusion
- MediaPipe selfie segmenter person mask, hands drawn in front by rule (no depth)
- Gestures
- Pinch state machine with hysteresis, arming and minimum durations; throw velocity is a least-squares fit before release
- Real-time budget
- One running + one replaceable pending job; stale answers discarded by snapshot id
What it doesn't do (yet)
- Segmentation varies with contrast, clutter, thin, transparent or reflective objects.
- Without a clean plate or scene memory, the hole is inpainted: expect a blurry patch. Shadows outside the mask stay visible.
- Image-plane rotation only. No depth, no 3D, no novel views; tracking is translation-only.
- The frames shown here are the real pipeline on the repo's synthetic desk fixture. A live webcam recording is still to be made.






