ICLR 2027 Submission Physical Evolution Supervision

EVEWorld

Physical Evolution Supervision for Embodied World Models

Process-level supervision for physically consistent target evolution.IGR restores target-instance consistency, while TIA promotes cross-frame target consistency.

#17 Overall WorldArena 2.0 · Track 1 #6 JEPA Similarity · EWMScore-P 70.17
−85.7% Model Laziness Rate 11.11% → 1.59% DreamGenBench · vs. matched Standard SFT
+13.6% Instruction Following Gemini-IF 53.57 → 60.85 · DreamGenBench
−26.5% Cross-backbone MLR 47.89% → 35.21% FlowWAM · RoboTwin
01Qualitative Comparison

Evolution supervision in action.

Matched rollouts under the same instruction compare Standard SFT, IGR, and the full EVEWorld model.

Standard SFT — frame-level reconstruction IGR — instance restoration supervision EVEWorld — restoration + temporal alignment
Case 01 / 05Qualitative rollout

InstructionUse the right hand to pick up red tomato from upper black tray of plastic shelf to inside of brown paper bag.

Standard SFT
IGROurs
EVEWorldOurs
Case 02 / 05Qualitative rollout

InstructionUse the right hand to pick up the bread slice from the toaster and place it on the blue plate.

Standard SFT
IGROurs
EVEWorldOurs
Case 03 / 05Qualitative rollout

InstructionMove the boxed drink the drink box with brown circular print onto the coaster the coaster made of smooth wood.

Standard SFT
IGROurs
EVEWorldOurs
Case 04 / 05Qualitative rollout

InstructionBring large block, medium block, and small block to the center using the left arm, the right arm, and the left arm.

Standard SFT
IGROurs
EVEWorldOurs
Case 05 / 05Qualitative rollout

InstructionPut the drink container with tapered shape squarely on top of the smooth round palm-sized coaster.

Standard SFT
IGROurs
EVEWorldOurs

Arrows indicate the conceptual addition of supervision components; variants use the same base initialization.

02Results & Generalization

WorldArena 2.0 submission and cross-setting generalization.

Quantitative and qualitative results are available in the paper; we summarize the strongest evidence here.

70.17 EWMScore-P Difficulty/OOD Adjusted
#6 JEPA Similarity
#17 Overall

Our FlowWAM-based EVEWorld submission, listed as Supervision_WM, achieves an EWMScore-P of 70.17 on the official WorldArena 2.0 Track 1 leaderboard, ranking 17th overall and 6th in JEPA Similarity.

Track 1 evaluates simulator video quality under WorldArena 2.0, with EWMScore-P reporting the difficulty- and out-of-distribution-adjusted aggregate score.

View Official Leaderboard ↗

WorldArena 2.0 Track 1 leaderboard showing the Supervision_WM submission with EWMScore-P 70.17 and overall rank 17.
Official WorldArena 2.0 Track 1 leaderboard. The FlowWAM-based EVEWorld submission is listed as Supervision_WM.

Protocol reference rollouts from the public evaluation harness — external baseline material rather than EVEWorld outputs — are collected on the protocol reference page.

DreamGenBench 11.11% → 1.59% Model Laziness Rate · −85.7% vs. matched Standard SFT
WorldArena 1.0 53.95 → 56.76 Overall · Dynamic Degree 48.49 → 62.90 · Flow 7.82 → 10.98 · Image Quality 60.71 → 59.82
EWMBench · AgiBot 3.7066 → 3.7525 Overall · Motion 61.51 → 63.65 · DYN 14.88 → 17.55 · Scene 84.38 → 86.37
FlowWAM · RoboTwin 47.89 → 35.21 MLR · PSNR 11.687 → 12.765 · SSIM 0.735 → 0.769 · LPIPS 0.414 → 0.365 · Flow-EPE 2.828 → 2.207

Reading these numbers. MLR is a dedicated target-instance consistency diagnostic rather than a composite measure of video quality, and it should be read together with motion, appearance, and instruction-following metrics. Not every component metric improves uniformly: IGR and TIA provide complementary supervision, and their relative contributions can depend on the backbone. Full profiles, the DreamGenBench splits, the component ablation, and the training setup are reported in the paper.

03Model Laziness

Two supervision gaps underlie inconsistent target evolution.

Standard frame-level reconstruction emphasizes local visual fidelity but provides limited process-level supervision over target evolution. This leaves two complementary gaps: target-instance consistency and cross-frame consistency.

Target-instance inconsistency

Model Laziness

Nt ≠ N0

Frame-level reconstruction provides weak supervision for maintaining the target-instance state throughout manipulation.

Unsupported duplicationUnsupported disappearance
Cross-frame inconsistency

Cross-frame consistency

Nt = N0 is necessary but not sufficient.

Correct instance count alone does not guarantee coherent target identity and appearance across time.

Identity driftDistortionFading
1

Missing instance-consistency supervision

The perturbed target region occupies only a small fraction of the full frame, allowing its reconstruction signal to be diluted by the surrounding scene. In controlled corruption experiments, Standard SFT retains 92–99% of the injected duplicate area.

IGR provides direct restoration supervision for the correct target-instance state.

2

Missing cross-frame consistency supervision

Correct target count alone does not establish correspondence between the target at adjacent frames. Consequently, locally plausible predictions may still drift in appearance, geometry, or identity over time.

TIA introduces explicit cross-frame alignment of target representations.

Comparison of Standard SFT, IGR, and EVEWorld: Standard SFT reaches the goal by duplicating the target, IGR suppresses duplication but leaves cross-frame distortion, EVEWorld preserves a single continuously evolving target
IGR substantially improves target-instance consistency, while residual cross-frame distortion can remain. Combining IGR with TIA promotes physically consistent target evolution across the interaction. (a) Standard SFT may reach the goal by duplicating the manipulated target. (b) IGR suppresses duplication, but cross-frame distortion can remain. (c) EVEWorld combines IGR and TIA to preserve a single target with continuous evolution throughout the interaction.
04Method

Instance restoration and temporal alignment for physically consistent target evolution.

EVEWorld retains the host model's native representation and introduces two complementary forms of evolution supervision during post-training: IGR promotes target-instance consistency through restoration supervision, while TIA promotes cross-frame consistency through temporal feature alignment.

Component 1

Instance-Guided Restoration (IGR)

Restoring the correct target-instance state from count-perturbed demonstrations.

  1. Construct count-edited demonstrations by inserting an additional target instance into clean training videos.
  2. Encode the clean and count-edited videos in the pretrained VAE latent space.
  3. Optimize restoration toward the clean latent representation, assigning 3× reconstruction weight to restoration-related regions and normalizing the spatial weight map to unit mean.

clean demo → + duplicate → count-edited demo → world model → clean target state

Component 2

Temporal Instance Alignment (TIA)

Aligning target representations across adjacent frames.

  1. Select the attachment layer using an offline layer-wise target-matching probe.
  2. Perform local cross-frame instance matching at the selected layer.
  3. Transport matched previous-frame features into the current representation through a bounded, zero-initialized residual, with correspondence supervision toward the true next-frame target location.

frame t−1 target → local matching → feature transport → ΔH → frame t representation

EVEWorld architecture: Instance-Guided Restoration on the left restores a clean video from a duplicate-corrupted input, Temporal Instance Alignment on the right aligns the target across adjacent frames
EVEWorld training framework. IGR constructs count-edited demonstrations and supervises restoration of the clean target state, while TIA performs cross-frame instance matching and residual feature transport at a selected transformer layer. At inference, EVEWorld requires only the initial image and instruction.

Training-only localization supervision. At inference, cross-frame matching operates directly on the model's hidden features and requires neither target localization nor an external tracker.

05Model Laziness Rate

Measuring persistent target-instance violations.

MLR measures persistent deviations from the initial target-instance count while excluding robot-supported occlusion.

Initial-state anchoredReference-freeOcclusion-awarePersistence-aware
D+Eligible set
D+ = { x ∈ D : N0(x) > 0 }

Only rollouts whose target is detected in the conditioning observation enter the denominator, so every model is evaluated on the same detector-eligible set.

MLRModel Laziness Rate
MLR = 100 / |D+|  ·  ∑x∈D+ 𝟙[ ∃t : Ñt(x) ≠ N0(x) ∧ Ñt+1(x) ≠ N0(x) ]

Computed per eligible video and reported as a percentage; lower is better. The adjusted count Ñt is the raw GroundingDINO count after robot-supported occlusions are discounted.

Eligibility is fixed from the conditioning observation. A rollout is counted as Model Laziness only when the adjusted target count remains inconsistent for at least two consecutive sampled timestamps.

MLR is a dedicated consistency diagnostic rather than a composite measure of overall video quality. It should therefore be interpreted jointly with motion, visual-quality, and instruction-following metrics. A low MLR alone does not imply strong dynamics or task execution.

How MLR counts: a red case where the deviation persists across two consecutive sampled frames is counted, a green case where a single frame is occluded is discarded
What counts as a violation. Red marks a count that stays deviated across two consecutive sampled frames and is therefore counted; green marks a single-frame deviation that the persistence rule discards.
06Qualitative Analysis

Target evolution across supervision variants.

Additional comparisons highlight instance consistency and cross-frame target consistency across representative manipulation tasks.

Pretrained — backbone before process-level supervision Standard SFT — frame-level reconstruction IGR — instance restoration supervision EVEWorld — restoration + temporal alignment

Supervision pathways

Starting from the same pretrained backbone, we compare Standard SFT, IGR, and the full EVEWorld model.

Case 01 / 02Supervision pathway

InstructionUse the right hand to pick up green apple from bottom shelf to top shelf.

Pretrained
Standard SFT
IGROurs
EVEWorldOurs
Case 02 / 02Supervision pathway

InstructionUse the left hand to pick up blue clock from bottom wooden shelf to center of teal plate.

Pretrained
Standard SFT
IGROurs
EVEWorldOurs

Connections illustrate supervision variants rather than sequential checkpoint initialization.

Direct SFT–EVEWorld comparisons

Representative side-by-side rollouts under matched instructions.

Case 01 / 04Matched comparison

InstructionPlace the white boxdrink with sturdy lid on the light gray streaked wooden coaster.

Standard SFT
EVEWorldOurs
Case 02 / 04Matched comparison

InstructionTake the bottle with green cap with arms, put it in the dustbin with white inner bag, repeat the same for the transparent bottle and the rounded bottle with a white neck.

Standard SFT
EVEWorldOurs
Case 03 / 04Matched comparison

InstructionGrab large block, medium block, small block, move them, and arrange by size centrally.

Standard SFT
EVEWorldOurs
Case 04 / 04Matched comparison

InstructionPlace the brown and white boxdrink on the wooden coaster.

Standard SFT
EVEWorldOurs
07Experiments

Four findings from our evaluation.

EVEWorld improves target-instance consistency across controlled ablations, backbone transfer, and stochastic inference.

01 · Main result

Evolution supervision sharply reduces Model Laziness.

−85.7%MLR

11.11% → 1.59%

DreamGenBench · vs. matched Standard SFT

Method MLR (%) ↓ Gemini-IF (%) ↑
GigaWorld-012.7060.19
Standard SFT11.1153.57
EVEWorld1.5960.85

Relative to the matched Standard-SFT baseline, EVEWorld reduces persistent target-instance violations while preserving instruction fidelity.

02 · Component ablation

IGR and TIA provide complementary supervision.

IGR + TIAboth components

1.59 MLR · 60.85 Gemini-IF

DreamGenBench · GigaWorld-based setting

Variant IGR TIA MLR (%) ↓ Gemini-IF (%) ↑
Standard SFT––11.1153.57
IGR only✓–4.7651.85
TIA only–✓7.9455.29
EVEWorld✓✓1.5960.85

IGR strengthens target-instance consistency, while TIA adds cross-frame supervision; combining both gives the strongest result on the GigaWorld-based setting.

03 · Cross-backbone

Evolution supervision transfers to FlowWAM.

47.89% → 35.21%MLR

Standard SFT → EVEWorld

FlowWAM · RoboTwin · held-out actions

Variant PSNR (dB) ↑ Flow-EPE ↓ MLR (%) ↓
Standard SFT11.6872.82847.89
+ IGR12.5442.48740.03
+ TIA13.4451.85022.54
EVEWorld12.7652.20735.21

TIA is particularly effective on this architecture, while the full EVEWorld model still improves substantially over Standard SFT, showing backbone-dependent component contributions.

04 · Robustness

Gains persist across stochastic inference trials.

6 / 6paired trials improve

PBench Robot QA

Gemini-3.8-Flash re-evaluation

3.8 s +1.90 +1.67 +1.04 Mean +1.54
5.8 s +0.60 +2.66 +1.42 Mean +1.56

Across three matched inference seeds and two rollout horizons, EVEWorld improves all six evaluated pairs.

Together, these results support evolution supervision across controlled ablations, backbone transfer, and stochastic inference.

08Citation

Cite this work

If you find this work useful, please cite it as follows.

@software{eveworld2026,
  title  = "{EVEWorld}: Physical Evolution Supervision for Embodied World Models",
  author = "{Anonymous Author(s)}",
  year   = {2026},
  url    = "{https://behappya.github.io/EVEWorld/}",
  note   = "Code and project page released for anonymous review."
}