Generating 6-DoF object trajectories by reading unrendered video diffusion latents — from a single image and one instruction.
1Southern University of Science and Technology 2Shanghai Innovation Institute 3Harbin Institute of Technology, Shenzhen 4SpatialtemporalAI
* Equal contribution † Corresponding author
An image and an instruction go into a frozen image-to-video diffusion model. At denoising step 16 of 40 we tap two signals out of it — a query–key attention track and pooled hidden states — and a small Reader turns them into a 6-DoF trajectory.
5,038 egocentric object trajectories across six corpora, each paired with a fine-grained natural-language instruction rather than a coarse verb–noun label.
The scene is reconstructed from the single input frame; the object is placed at the 13 predicted poses. Drag to orbit, right-drag to pan, scroll to zoom — the camera is unconstrained, so it can leave the region the input image covers. Scrub the timeline to step through the trajectory, and press ⟲ to return to the original viewpoint.