MoT world model + wrist cameras — React motherboard

every image stacks view / wrist L / wrist R / tactile L / tactile R, each as a GT row over a PRED row · red box = first predicted frame · 15 fps · sections: React motherboard (v3mb)

React motherboard (v3mb)

run mot_df_vt_a9dfilm_dforce_isolate_adaln_wrist_v3mb

Trained on React motherboard 2026-09-11/09-12 (25 segments, 12 recordings, 84,107 windows). val = test (the split aliases them): the separate 2026-09-09 session (motherboard episode_000-002, 9,488 windows) — different wrist cameras (neutral grey vs the warm 09-11 cameras) and a different calibration epoch, so it measures domain shift, not a random held-out split; the model repaints the 09-09 wrist cameras into the training-camera look and scores below the copy-last-frame floor on every stream. train = 4 training episodes (09-11 ep000, ep005_seg04; 09-12 ep000_seg00, ep003). All sections use last.ckpt pinned at epoch 57 (training was still running).
checkpoint:

Long horizon ( s, keep-1 autoregressive) —

mean PSNR vs time — — dotted line = end of GT context
video: GT | PRED for all five streams
long rollout strip

Contact-stratified short horizon —

Windows of each split classed by the recorded GelSight normal force (max of L/R, contact = > 0.5 N) over the 16 frames: none = no contact, full = contact in ≥ 80 % of frames, onset = contact starts inside the window, release = contact ends inside it. 125 windows per class and split, one sample each. true = model with the recorded actions, floor = copy the last context frame, static = model with frozen actions, moving = PSNR inside the pixels that move in GT.
contact-class strip

Short horizon (2 context + 2 predicted latents) —

short rollout strip

Training progression (4 fixed train windows, every epoch, 20 sampling steps)

val loss
min val loss
training strip