Skip to content

AMAP-ML/DreamX-World

Repository files navigation

DreamX-World: A General-Purpose Interactive World Model

DreamX Team

Official Website Reactor HuggingFace ModelScope Tech Report

Page License


DreamX-World is a general-purpose world model for interactive world simulation. It generates diverse, high-fidelity worlds that users can explore, control, and transform with event prompts.

The model is trained with a scalable data engine on Unreal Engine data, gameplay footage, and real-world videos, combined with camera estimation and strict data filtering to learn realistic dynamics and interactions. It follows a progressive training pipeline: learning fine-grained action control first, then open-ended event response, and using Reinforcement Learning to improve action following, interaction consistency, and visual fidelity. Finally, through forcing and distillation, DreamX-World achieves efficient inference, making interactive generation practical at scale.

🔥 News

  • Jul 23, 2026: DreamX-World 1.0 is now live! Head over to the Official Website to experience our interactive world model.
  • June 15, 2026: We released DreamX-World 1.0 technical report.
  • June 15, 2026: We open-sourced DreamX-World-5B that supports 1-min video generation.
  • May 11, 2026: We open-sourced DreamX-World-5B-Cam and inference codes.

📆 Plan

  • ✔️ DreamX-World-5B-Cam Model.
  • ✔️ Long-horizon DreamX-World-5B Model.
  • ✔️ Release Technical Report.
  • DreamX-World-14B-Cam Model.
  • Audio-Video Joint Generation Model.

🚀 Quick Start

Setup

  1. Install dependencies
pip install -r requirements.txt
  1. Download Wan2.2-5B-TI2V checkpoints from https://huggingface.co/Wan-AI

Inference

Please check out inference_README.md for detailed instructions.

📍 Checkpoints

Model Download Link Details Instrutions
DreamX-World-5B-Cam Huggingface, ModelScope Bidrectional, Supports 5s Video Generation inference_README.md
DreamX-World-5B Huggingface, ModelScope Autoregressive, Supports Long-horizon Video Generation inference_README.md

Computational Efficiency

The inference time (shown in the Cost column below) comprises both denoising and VAE decoding time.

Model GPUs Video Cost (Time in second/Peak Memory)
DreamX-World-5B-Cam 1xH20 5s 720P 509s/38G
DreamX-World-5B 1xH20 5s 720P 26s/40G
DreamX-World-5B 1xH20 60s 720P 342s/72G

🎬 Video Demo

Note: The demo videos are intentionally compressed to ensure smooth playback, which may result in a slight loss of visual quality.

⏳ Generate Long-Horizon Worlds

DreamX-World supports long-horizon autoregressive generation with precise camera control. Progressive training on long rollouts mitigates identity, background, style, and color drift, enabling coherent world exploration over hundreds of frames.

long_01.mp4
long_05.mp4
long_03.mp4
long_04.mp4

🧠 Remember and Revisit

DreamX-World uses geometry-guided memory retrieval to recover non-local visual evidence from earlier observations. This improves scene persistence when the camera revisits a previously explored region, preserving its layout, object identities, and local appearance.

memory_01.mp4
memory_02.mp4
memory_03.mp4
memory_04.mp4
memory_05.mp4

🌍 Navigate and Explore Realistic Worlds

DreamX-World enables high-fidelity, controllable exploration across diverse realistic environments, including indoor, urban, natural, and architectural scenes.

01.mp4
02.mp4
03.mp4
04.mp4
05.mp4
06.mp4
07.mp4
08.mp4

🌈 Dive into Dream Worlds

Beyond realistic scenes, DreamX-World also generates fantasy, game-like, sci-fi, and stylized worlds.

01.mp4
02.mp4
03.mp4
04.mp4
06.mp4
07.mp4
08.mp4
09.mp4

🎮 Generate in Third-Person View

DreamX-World supports both first-person interaction and coherent third-person generation. It keeps camera-follow behavior stable while preserving controllable agent motion and scene consistency.

01.mp4
02.mp4
03.mp4
04.mp4
05.mp4
07.mp4
08.mp4
10.mp4

⚡ Promptable World Events

DreamX-World supports prompt-driven world events that dynamically change the environment, including flexible and compositional event generation with consistent temporal evolution.

  • Single Event: A single event prompt triggers a specific world-changing interaction.
  • Compositional Events: Multiple events compose together to create complex, multi-step world transformations.

Single Event

01.mp4
02.mp4
03.mp4
04.mp4

Compositional Events

05.mp4
06.mp4
07.mp4
08.mp4

💬 WeChat Group

Join our WeChat group for discussion:

WeChat Group QR Code

📚 Citation

If you find DreamX-World useful in your research, please consider citing our technical report:

@article{team2026dreamx,
  title={DreamX-World 1.0: A General-Purpose Interactive World Model},
  author={Team, DreamX and Bai, Yancheng and Chen, Rui and Chu, Xiangxiang and Dang, Rujing and Dou, Hao and Gao, Bingjie and Gu, Qiwen and Hong, Siyu and Lei, Jiachen and others},
  journal={arXiv preprint arXiv:2606.16993},
  year={2026}
}

📜 License

This project is licensed under Apache 2.0. See LICENSE for details.

✨ Acknowledgement

We thank the Wan Team for open-sourcing their code and models.

About

DreamX-World: A General-Purpose Interactive World Model

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages