DreamTraj

Generating 6-DoF object trajectories by reading unrendered video diffusion latents — from a single image and one instruction.

Tongsheng Ding1 *Zhen Luo2,1 * Yixuan Yang1 *Boyu Wang1 Luyang Xie1Jinyu Yang3 † Feng Zheng1,4

1Southern University of Science and Technology  2Shanghai Innovation Institute  3Harbin Institute of Technology, Shenzhen  4SpatialtemporalAI

* Equal contribution   † Corresponding author

Method

An image and an instruction go into a frozen image-to-video diffusion model. At denoising step 16 of 40 we tap two signals out of it — a query–key attention track and pooled hidden states — and a small Reader turns them into a 6-DoF trajectory.

DreamTraj framework overview

MOVE dataset

5,038 egocentric object trajectories across six corpora, each paired with a fine-grained natural-language instruction rather than a coarse verb–noun label.

Interactive results

The scene is reconstructed from the single input frame; the object is placed at the 13 predicted poses. Drag to orbit, right-drag to pan, scroll to zoom — the camera is unconstrained, so it can leave the region the input image covers. Scrub the timeline to step through the trajectory, and press ⟲ to return to the original viewpoint.

the single input frame this scene was reconstructed from anchor frame — the only image the model sees · click to hide
0 / 12