Telekinesis

Can you pick up anything on your desk, in video, without a class list or a 3D model?

Touch a real object in the webcam image, pinch, and lift its appearance out of the video while the background is reconstructed behind it.

0.3–1.1 sEdgeSAM encode on a laptop CPU, run once per selection; tracking is per-frame
Year
2026
Status
Working local app · CPU only · honest limitations documented
Stack
Python, OpenCV, MediaPipe, EdgeSAM, ONNX Runtime, NumPy
01

One interaction, eight frames

These frames come from running the project's own fixture script: real EdgeSAM, real app logic, a drawn hand driving synthetic landmarks over a synthetic desk. Nothing is mocked except the camera.

Step 1: Raw scene. Nothing selected.

Real pipeline output on the repo's synthetic desk fixture. No webcam footage.

02

Detection, segmentation and tracking are different jobs

Detection

MediaPipe finds hands. No object detector and no class list: anything coherent can be picked up.

Segmentation

EdgeSAM, a promptable SAM-family model, answers “what object is at this point?” once per selection, in a background thread. The prompt sits slightly ahead of the fingertip, because the fingertip pixel is skin; points on the hand go in as negative prompts.

Choosing the whole object

SAM's own score favours small crisp parts. Ranking drops implausible masks, then prefers the largest candidate whose score and stability (does the mask change if the logit threshold moves ±1?) are close to the best. Stability is what rejects merges of neighbouring objects.

Tracking

Masked normalised cross-correlation in a small window follows the same object every frame. It can report OCCLUDED or LOST and hold still; it never switches objects.

One more section for engineers: architecture, hyper-parameters and design decisions.

03

Under the hood

Under the hood
Transforms
BGR→RGB, SAM normalisation, longest side 1024, pad; logits upsampled, padding cropped, then resized back
Compositing
Feathered soft alpha; premultiplied colour before warping to avoid dark fringes
Geometry
One affine T(pos)·R(angle)·S(scale)·T(−anchor) applied to pixels and alpha alike
Reconstruction
Clean plate → scene memory (last ~40 s) → pyramid fill, with OpenCV FSR in the background
Occlusion
MediaPipe selfie segmenter person mask, hands drawn in front by rule (no depth)
Gestures
Pinch state machine with hysteresis, arming and minimum durations; throw velocity is a least-squares fit before release
Real-time budget
One running + one replaceable pending job; stale answers discarded by snapshot id
04

What it doesn't do (yet)

  1. Segmentation varies with contrast, clutter, thin, transparent or reflective objects.
  2. Without a clean plate or scene memory, the hole is inpainted: expect a blurry patch. Shadows outside the mask stay visible.
  3. Image-plane rotation only. No depth, no 3D, no novel views; tracking is translation-only.
  4. The frames shown here are the real pipeline on the repo's synthetic desk fixture. A live webcam recording is still to be made.

Read the code and the full results on GitHub