Preprint · World Action Models

FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models

1 Tsinghua University

2 TeleAI, China Telecom

3 Shenzhen Technology University

† Corresponding authors. * Work done during an internship at TeleAI.

Comparison between a unified future objective and FutureDuet's view-specific future modeling
Two views for control. Two predictive roles for learning.

01 · Insight

One policy, two predictive roles

Stable scene frame

Main view → Task State

Model scene-level progress, with masks and skeletons focusing supervision on task objects and robot motion.

Moving camera frame

Wrist views → Interaction State

Model short-horizon interaction changes without reconstructing every camera-induced appearance change.

Use both views for control; supervise each view according to the future it reveals.

02 · Method

Task evolution meets interaction evolution

View-specific objectives shape two compact states; the Duet Interface reunites them for action generation.

FutureDuet architecture with main-view and wrist-view streams joining ActionDiT
Future prediction shapes representation learning during training; deployment retains the compact control path.

03 · Results

Broad gains, strongest on precise interaction

The largest RoboTwin50 gains appear where success depends on precise contact geometry.

94.2%

RoboTwin50
clean

94.1%

RoboTwin50
randomized

99.2%

LIBERO
average

+13.9 pp

Separate future latent
over unified RGB

Wrist pathway study

Separation helps; temporal supervision helps further

With the deployed route fixed, future RGB improves success to 64.1% and future latent supervision to 69.7%, at the same 201.20 ms latency.

Success and latency trade-off for wrist pathway designs
RoboTwin50 task threshold counts and gains on challenging tasks
FutureDuet raises the long tail, with its largest gains on demanding interaction tasks.

Where difficulty remains

Reaching the object is not the same as completing the interaction

Residual failures arise at the final opening, hooking, or contact stage.

Failure cases for Open Microwave, Hanging Mug, and Turn Switch

04 · Readout evidence

Structured targets shape what actions read

Masks emphasize task objects; skeletons emphasize robot motion.

Qualitative comparison of action readout under RGB-only and structured supervision
Heatmap of target-specific changes in object and robot readout

05 · Citation

Cite FutureDuet

@misc{wu2026futureduet,
  title  = {FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models},
  author = {Jie Wu and Yuzhi Huang and Junqi Liu and Weichen Zhang and Haibin Huang and Yin Chen and Jingyan Jiang and Chi Zhang},
  year   = {2026},
  eprint = {2609.34362},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url    = {https://arxiv.org/abs/2609.34362}
}