InstructionUse the right hand to pick up red tomato from upper black tray of plastic shelf to inside of brown paper bag.
Standard SFT
IGROurs
EVEWorldOurs
Case 02 / 05Qualitative rollout
InstructionUse the right hand to pick up the bread slice from the toaster and place it on the blue plate.
Standard SFT
IGROurs
EVEWorldOurs
Case 03 / 05Qualitative rollout
InstructionMove the boxed drink the drink box with brown circular print onto the coaster the coaster made of smooth wood.
Standard SFT
IGROurs
EVEWorldOurs
Case 04 / 05Qualitative rollout
InstructionBring large block, medium block, and small block to the center using the left arm, the right arm, and the left arm.
Standard SFT
IGROurs
EVEWorldOurs
Case 05 / 05Qualitative rollout
InstructionPut the drink container with tapered shape squarely on top of the smooth round palm-sized coaster.
Standard SFT
IGROurs
EVEWorldOurs
Arrows indicate the conceptual addition of supervision components; variants use the same base initialization.
02Results & Generalization
WorldArena 2.0 submission and cross-setting generalization.
Quantitative and qualitative results are available in the paper; we summarize the strongest evidence here.
70.17EWMScore-PDifficulty/OOD Adjusted
#6JEPA Similarity
#17Overall
Our FlowWAM-based EVEWorld submission, listed as Supervision_WM, achieves an EWMScore-P of 70.17 on the official WorldArena 2.0 Track 1 leaderboard, ranking 17th overall and 6th in JEPA Similarity.
Track 1 evaluates simulator video quality under WorldArena 2.0, with EWMScore-P reporting the difficulty- and out-of-distribution-adjusted aggregate score.
Official WorldArena 2.0 Track 1 leaderboard. The FlowWAM-based EVEWorld submission is listed as Supervision_WM.
Protocol reference rollouts from the public evaluation harness — external baseline material rather than EVEWorld outputs — are collected on the protocol reference page.
DreamGenBench11.11% → 1.59%Model Laziness Rate · −85.7% vs. matched Standard SFT
Reading these numbers. MLR is a dedicated target-instance consistency diagnostic rather than a composite measure of video quality, and it should be read together with motion, appearance, and instruction-following metrics. Not every component metric improves uniformly: IGR and TIA provide complementary supervision, and their relative contributions can depend on the backbone. Full profiles, the DreamGenBench splits, the component ablation, and the training setup are reported in the paper.
03Model Laziness
Two supervision gaps underlie inconsistent target evolution.
Standard frame-level reconstruction emphasizes local visual fidelity but provides limited process-level supervision over target evolution. This leaves two complementary gaps: target-instance consistency and cross-frame consistency.
Target-instance inconsistency
Model Laziness
Nt ≠ N0
Frame-level reconstruction provides weak supervision for maintaining the target-instance state throughout manipulation.
Unsupported duplicationUnsupported disappearance
Cross-frame inconsistency
Cross-frame consistency
Nt = N0 is necessary but not sufficient.
Correct instance count alone does not guarantee coherent target identity and appearance across time.
Identity driftDistortionFading
1
Missing instance-consistency supervision
The perturbed target region occupies only a small fraction of the full frame, allowing its reconstruction signal to be diluted by the surrounding scene. In controlled corruption experiments, Standard SFT retains 92–99% of the injected duplicate area.
IGR provides direct restoration supervision for the correct target-instance state.
2
Missing cross-frame consistency supervision
Correct target count alone does not establish correspondence between the target at adjacent frames. Consequently, locally plausible predictions may still drift in appearance, geometry, or identity over time.
TIA introduces explicit cross-frame alignment of target representations.
IGR substantially improves target-instance consistency, while residual cross-frame distortion can remain. Combining IGR with TIA promotes physically consistent target evolution across the interaction. (a) Standard SFT may reach the goal by duplicating the manipulated target. (b) IGR suppresses duplication, but cross-frame distortion can remain. (c) EVEWorld combines IGR and TIA to preserve a single target with continuous evolution throughout the interaction.
04Method
Instance restoration and temporal alignment for physically consistent target evolution.
EVEWorld retains the host model's native representation and introduces two complementary forms of evolution supervision during post-training: IGR promotes target-instance consistency through restoration supervision, while TIA promotes cross-frame consistency through temporal feature alignment.
Component 1
Instance-Guided Restoration (IGR)
Restoring the correct target-instance state from count-perturbed demonstrations.
Construct count-edited demonstrations by inserting an additional target instance into clean training videos.
Encode the clean and count-edited videos in the pretrained VAE latent space.
Optimize restoration toward the clean latent representation, assigning 3× reconstruction weight to restoration-related regions and normalizing the spatial weight map to unit mean.
clean demo → + duplicate → count-edited demo → world model → clean target state
Component 2
Temporal Instance Alignment (TIA)
Aligning target representations across adjacent frames.
Select the attachment layer using an offline layer-wise target-matching probe.
Perform local cross-frame instance matching at the selected layer.
Transport matched previous-frame features into the current representation through a bounded, zero-initialized residual, with correspondence supervision toward the true next-frame target location.
frame t−1 target → local matching → feature transport → ΔH → frame t representation
EVEWorld training framework. IGR constructs count-edited demonstrations and supervises restoration of the clean target state, while TIA performs cross-frame instance matching and residual feature transport at a selected transformer layer. At inference, EVEWorld requires only the initial image and instruction.
Training-only localization supervision. At inference, cross-frame matching operates directly on the model's hidden features and requires neither target localization nor an external tracker.
05Model Laziness Rate
Measuring persistent target-instance violations.
MLR measures persistent deviations from the initial target-instance count while excluding robot-supported occlusion.
Only rollouts whose target is detected in the conditioning observation enter the denominator, so every model is evaluated on the same detector-eligible set.
Computed per eligible video and reported as a percentage; lower is better. The adjusted count Ñt is the raw GroundingDINO count after robot-supported occlusions are discounted.
Eligibility is fixed from the conditioning observation. A rollout is counted as Model Laziness only when the adjusted target count remains inconsistent for at least two consecutive sampled timestamps.
MLR is a dedicated consistency diagnostic rather than a composite measure of overall video quality. It should therefore be interpreted jointly with motion, visual-quality, and instruction-following metrics. A low MLR alone does not imply strong dynamics or task execution.
What counts as a violation.Red marks a count that stays deviated across two consecutive sampled frames and is therefore counted; green marks a single-frame deviation that the persistence rule discards.
06Qualitative Analysis
Target evolution across supervision variants.
Additional comparisons highlight instance consistency and cross-frame target consistency across representative manipulation tasks.
Starting from the same pretrained backbone, we compare Standard SFT, IGR, and the full EVEWorld model.
Case 01 / 02Supervision pathway
InstructionUse the right hand to pick up green apple from bottom shelf to top shelf.
Pretrained
Standard SFT+ IGR
Standard SFT
IGROurs
EVEWorldOurs
Case 02 / 02Supervision pathway
InstructionUse the left hand to pick up blue clock from bottom wooden shelf to center of teal plate.
Pretrained
Standard SFT+ IGR
Standard SFT
IGROurs
EVEWorldOurs
Connections illustrate supervision variants rather than sequential checkpoint initialization.
Direct SFT–EVEWorld comparisons
Representative side-by-side rollouts under matched instructions.
Case 01 / 04Matched comparison
InstructionPlace the white boxdrink with sturdy lid on the light gray streaked wooden coaster.
Standard SFT
EVEWorldOurs
Case 02 / 04Matched comparison
InstructionTake the bottle with green cap with arms, put it in the dustbin with white inner bag, repeat the same for the transparent bottle and the rounded bottle with a white neck.
Standard SFT
EVEWorldOurs
Case 03 / 04Matched comparison
InstructionGrab large block, medium block, small block, move them, and arrange by size centrally.
Standard SFT
EVEWorldOurs
Case 04 / 04Matched comparison
InstructionPlace the brown and white boxdrink on the wooden coaster.
Standard SFT
EVEWorldOurs
07Experiments
Four findings from our evaluation.
EVEWorld improves target-instance consistency across controlled ablations, backbone transfer, and stochastic inference.
01 · Main result
Evolution supervision sharply reduces Model Laziness.
−85.7%MLR
11.11% → 1.59%
DreamGenBench · vs. matched Standard SFT
Method
MLR (%) ↓
Gemini-IF (%) ↑
GigaWorld-0
12.70
60.19
Standard SFT
11.11
53.57
EVEWorld
1.59
60.85
Relative to the matched Standard-SFT baseline, EVEWorld reduces persistent target-instance violations while preserving instruction fidelity.
02 · Component ablation
IGR and TIA provide complementary supervision.
IGR + TIAboth components
1.59 MLR · 60.85 Gemini-IF
DreamGenBench · GigaWorld-based setting
Variant
IGR
TIA
MLR (%) ↓
Gemini-IF (%) ↑
Standard SFT
–
–
11.11
53.57
IGR only
✓
–
4.76
51.85
TIA only
–
✓
7.94
55.29
EVEWorld
✓
✓
1.59
60.85
IGR strengthens target-instance consistency, while TIA adds cross-frame supervision; combining both gives the strongest result on the GigaWorld-based setting.
03 · Cross-backbone
Evolution supervision transfers to FlowWAM.
47.89% → 35.21%MLR
Standard SFT → EVEWorld
FlowWAM · RoboTwin · held-out actions
Variant
PSNR (dB) ↑
Flow-EPE ↓
MLR (%) ↓
Standard SFT
11.687
2.828
47.89
+ IGR
12.544
2.487
40.03
+ TIA
13.445
1.850
22.54
EVEWorld
12.765
2.207
35.21
TIA is particularly effective on this architecture, while the full EVEWorld model still improves substantially over Standard SFT, showing backbone-dependent component contributions.
04 · Robustness
Gains persist across stochastic inference trials.
6 / 6paired trials improve
PBench Robot QA
Gemini-3.8-Flash re-evaluation
3.8 s+1.90+1.67+1.04Mean +1.54
5.8 s+0.60+2.66+1.42Mean +1.56
Across three matched inference seeds and two rollout horizons, EVEWorld improves all six evaluated pairs.
Together, these results support evolution supervision across controlled ablations, backbone transfer, and stochastic inference.
08Citation
Cite this work
If you find this work useful, please cite it as follows.
@software{eveworld2026,
title = "{EVEWorld}: Physical Evolution Supervision for Embodied World Models",
author = "{Anonymous Author(s)}",
year = {2026},
url = "{https://behappya.github.io/EVEWorld/}",
note = "Code and project page released for anonymous review."
}