What Makes World Action Models Generalize?
An Empirical Study of Test-Time Future Modeling

1Leap Lab, Tsinghua University, 2University of Science and Technology of China, 3Beijing Institute of Technology
*Equal contribution ✉Corresponding author

TL;DR: Future conditioning, rather than clean future generation, is what makes World Action Models generalize. Simple-WAM uses a single pass over fully noised future tokens, achieving stronger generalization across simulation and real-world tasks while retaining efficiency comparable to Latent WAMs.

Simple-WAM generalization study and its two main findings

What makes World Action Models generalize? Explicit WAMs repeatedly denoise future frames at inference, while Latent WAMs remove future modeling for efficiency. We find that future conditioning is essential for generalization, but clean future generation is not: a single pass over fully noised future tokens is already sufficient. Based on this finding, Simple-WAM retains the generalization benefit of explicit future modeling at a cost close to latent inference.

1 Abstract

World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirely for acceleration.

We find that latent WAMs, despite matching explicit ones on in-distribution tasks, fail to retain the generalization benefits that originally motivated WAMs. To demonstrate this, we evaluate generalization along three axes: environmental perturbation, data efficiency, and task generalization. Controlled comparisons with a matched backbone, training data, and budget reveal consistent degradation across all three axes when the action expert no longer conditions on future representations. Further analysis shows that the gap arises almost entirely from the first denoising step: the benefit comes from preparing the future, not generating it.

We therefore propose Simple-WAM, which simplifies future modeling into a single forward pass of fully noised video tokens and adapts the training-time noise schedule to this inference behavior. Across simulation and real-world tasks, Simple-WAM achieves the best of both worlds, leading explicit WAMs in generalization performance with efficiency comparable to Latent WAMs.

2 Empirical Findings

Finding 2.1

Generalization Deteriorates without Inference-Time Future Modeling

Explicit and latent WAM success rates under in-distribution and three generalization settings

Explicit and latent WAMs are within 0.9 points on in-distribution tasks, yet they separate consistently under environmental perturbations, reduced demonstrations, and held-out tasks. The latent policy loses 14.0 points, 8.5 points, and 34.1 points on the three axes, respectively. Test-time future conditioning is therefore an important source of generalizable temporal information.

Finding 2.2

A Single Pass over a Fully Noised Future Is Sufficient

Generalization performance across future-video denoising steps

Most of the generalization benefit appears at the first video-model pass. Future video tokens can remain at pure Gaussian noise: the action expert needs the future-conditioned feature, rather than a decoded clean prediction of the future.

3 Method

Comparison of Explicit WAM, Latent WAM, and Simple-WAM

One pass over a fully noised future. At inference, Simple-WAM keeps future video tokens at Gaussian noise and runs the video expert once. The resulting future-conditioned feature is reused by every action denoising step.

A mixed training schedule. Training samples the fully noised endpoint with probability p = 0.5 while retaining the original flow-time schedule for the remaining samples. This aligns training with the single flow time used at inference without adding modules or objectives.

4 Experiments

We evaluate Simple-WAM along environmental perturbation, data efficiency, and task generalization in simulation and on a real ALOHA dual-arm robot.

4.1 Overview

LIBERO-Plus

Seven perturbation factors: robot initialization, camera viewpoint, language, sensor noise, background, object layout, and lighting.

LIBERO

Data efficiency at 5 and 10 demonstrations per task, and task generalization with and without action-free video.

RoboTwin 2.0

Bimanual manipulation under 10-shot learning and held-out tasks, with held-out settings matching LIBERO.

Real-world robot tasks and evaluation conditions
Real-world setup. Four training tasks, environmental perturbations, and the held-out Store in Order task on an ALOHA dual-arm platform.

4.2 Quantitative Results

Method Environmental Perturbation Data Efficiency Task Generalization
LIBERO-PlusAverage LIBERO5-Shot LIBERO10-Shot RoboTwin10-Shot LIBEROw/o Video LIBEROw/ Video RoboTwinw/o Video RoboTwinw/ Video
FastWAM 53.8 78.288.54.8 2.15.93.54.8
FastWAM-Joint 67.7 91.597.035.0 6.369.95.447.3
Simple-WAM 79.5 92.497.237.2 10.173.66.244.5
Simulation results. Success rates (%) on LIBERO-Plus, LIBERO, and RoboTwin 2.0. Bold values indicate the best result in each setting.
Real-world robot evaluation results arranged in a two-by-two grid
Real-world results. Success rate across in-distribution tasks, environmental perturbations, data efficiency, and task generalization.
Inference latency and generalization performance
Efficiency. Per-chunk latency and average success rate over four generalization settings. Simple-WAM runs at 74.7 ms per action chunk, 3.8× faster than FastWAM-Joint.

4.3 Qualitative Results

Environmental Perturbation

Robot Initial State

Example 1
Example 2
Example 3

Camera Viewpoint

Example 1
Example 2
Example 3

Language Instruction

Prompt 1“Could you move the black bowl from the stove to the plate?”
Prompt 2“Lift the darkcolored rounded container situated beside the container holding baked sweet treats and set it down upon the flat circular dish.”
Prompt 3“Put the black bowl from the wooden cabinet onto the plate.”

Sensor Noise

Example 1
Example 2
Example 3

Background Texture

Example 1
Example 2
Example 3

Object Layout

Example 1
Example 2
Example 3

Lighting Condition

Example 1
Example 2
Example 3
Task 1Stack the Bowl Task 2Move Forward

Data Efficiency

Open the Middle Drawer
Alphabet Soup to Basket
Black Bowl to Plate
Adjust Bottle
Hang the Mug
Open Laptop
✓Success ×Fail ✓Success ✓Success

Task Generalization

Put the Bowl on the Stove

×Fail
Without Action-Free Video
✓Success
With Action-Free Video

Alphabet Soup to Basket

×Fail
Without Action-Free Video
✓Success
With Action-Free Video

Black Bowl to Plate

✓Success
Without Action-Free Video
✓Success
With Action-Free Video

Beat Block with Hammer

Without Action-Free Video
With Action-Free Video

Blocks Ranking

Without Action-Free Video
With Action-Free Video

Place Cans in Plastic Box

Without Action-Free Video
With Action-Free Video

Unseen TaskPick Up the Object and Place It onto the Plate

Prompt“Pick up the object and place it onto the plate. Follow this order: red, yellow, green.” Prompt“Pick up the object and place it onto the plate. Follow this order: yellow, green, red.”

BibTeX

@article{zhou2026simplewam,
  title={What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling},
  author={Zhou, Renping and Ni, Zanlin and Fan, Zihao and Fu, Guohao and Liu, Zeyu and Shi, Hao and Zhang, Jie and Chen, Chi Bene and Yue, Yang and Fu, Xueyang and Huang, Gao},
  journal={arXiv preprint},
  year={2026}
}