MoT world model + wrist cameras — React motherboard
every image stacks view / wrist L / wrist R / tactile L / tactile R, each as a GT row over a PRED row · red box = first predicted frame · 15 fps · sections: React motherboard (v3mb)
React motherboard (v3mb)
run mot_df_vt_a9dfilm_dforce_isolate_adaln_wrist_v3mb
Trained on React motherboard 2026-09-11/09-12 (25 segments, 12 recordings, 84,107 windows). val = test (the split aliases them): the separate 2026-09-09 session (motherboard episode_000-002, 9,488 windows) — different wrist cameras (neutral grey vs the warm 09-11 cameras) and a different calibration epoch, so it measures domain shift, not a random held-out split; the model repaints the 09-09 wrist cameras into the training-camera look and scores below the copy-last-frame floor on every stream. train = 4 training episodes (09-11 ep000, ep005_seg04; 09-12 ep000_seg00, ep003). All sections use last.ckpt pinned at epoch 57 (training was still running).
checkpoint:
Long horizon ( s, keep-1 autoregressive) —
mean PSNR vs time — — dotted line = end of GT context
video: GT | PRED for all five streams
Contact-stratified short horizon —
Windows of each split classed by the recorded GelSight normal force (max of L/R, contact = > 0.5 N) over the 16 frames:
none = no contact, full = contact in ≥ 80 % of frames, onset = contact starts inside the window, release = contact ends inside it.
125 windows per class and split, one sample each. true = model with the recorded actions, floor = copy the last context frame,
static = model with frozen actions, moving = PSNR inside the pixels that move in GT.
Short horizon (2 context + 2 predicted latents) —
Training progression (4 fixed train windows, every epoch, 20 sampling steps)
val loss
min val loss