Local AI World Model Part 2 - Deep NN to turn Images into Playable Characters, with prompt switching mid rollout
The idea was simple :
Can a neural network generate an interactive world in real time, locally, on consumer hardware?
That older model could take an image and my keyboard inputs and autoregressively generate the next frames while I played with it...
But it had one major limitation:
it basically couldn't understand text.
So for the past 2 months, I've been training a new model.
This time, the model can be controlled with both keyboard actions + live text prompts.
The Demo
The video above starts from a single image of a man in a desert, I got it off google images.
While the model is generating the world, I can change the prompt live.
For example:
"add a pond"
or
"put the man in a hoodie"
and the model tries to modify the world while continuing the same rollout.
At the same time, it is still taking action inputs and generating the next frames autoregressively.
Everything is happening live.
The demo is running on an RTX 5090, but the model is intentionally software-throttled to around 12 FPS.
It is a roughly 960M parameter model, and depending on the configuration I can get around 50–60 generated frames/sec on a 5090 peak fps.
The GPU in the demo is often sitting at only around 30% utilization.
I haven't properly benchmarked the 30-series and 40-series cards yet, so I don't want to make hard claims there until I actually test them, and will do that soon.
Architecture
The model is a pure transformer with 28 blocks and 20 attention heads.
It uses a block-causal attention mask.
Conceptually, it behaves like an LLM. During training, earlier frames are prevented from seeing future frames. During inference, once a frame has finished denoising, its attention keys and values are added to the KV cache. Then the model generates the next frame.
So instead of repeatedly processing the entire video history, the model can reuse information from previous frames.
The main difference from an LLM is, We don't keep the entire history. The KV cache uses a sliding context window, currently corresponding to roughly 80 latent frames. Otherwise the cost of attention would just keep increasing forever.
Diffusion Forcing
Another important part of the training setup is something called diffusion forcing. It is a brilliant paper. In normal diffusion video training, frames are often corrupted together according to the same diffusion timestep. Instead, we independently noise different frames during training. So one frame might be almost clean while another frame is extremely noisy.
The reason is that autoregressive generation can easily get noisy and turn into garbage. At inference time, the model will inevitably have to condition on its own imperfect previous generations.
Training with independently corrupted history makes it much more comfortable operating when its past context isn't perfect.
Why the Previous Model Couldn't Do This
The previous model used an MMDiT-style architecture. That meant text and video information were mixed through the same attention mechanism. For a normal text-to-video model this can work perfectly well.
But for an autoregressive world model, the amount of visual context becomes enormous. As the rollout grows, you can have thousands and thousands of visual KV tokens competing with a relatively tiny amount of text conditioning. That made reliable live prompt changes extremely difficult.
So the new model goes back to explicit text cross-attention. The visual history remains in the causal self-attention KV cache, while text is provided separately through cross-attention. I also did substantially more text-video pretraining this time.
That combination made the difference much larger than I expected.
Action Control
The model also receives the keyboard inputs directly. For now these are simple controls like: W A S D The action signal is injected into the transformer through an additional conditioning term in the adaLN, so its a simple addition to the timestep emebdding.
The long-term goal is obviously much broader action spaces than four keyboard keys, but this is a very convenient environment for testing whether the basic idea works.
Training
The model was trained using 8× H100 SXM GPUs for roughly 3–4 weeks. Unlike a lot of recent realtime video systems, this is not an autoregressive distillation of WAN, LTX, or another large pretrained video generator.
The model was trained specifically around this architecture and realtime inference constraint. That constraint is actually very important to me. I want local inference to be a design constraint from the beginning.
Every model I make for this project will be built around that.
MacBook
I'm also very interested in getting this running outside NVIDIA GPUs. I haven't finished the MLX port yet. However, I wrote a benchmark for the core transformer computation and I'm seeing roughly 30 FPS-equivalent throughput on my M5 MacBook. The problem is the autoencoder, and CNNs have very poor kernels on macbooks so the forward pass is 100-200ms instead of the 10ms you get on RTX GPUs.
But the result made me considerably more optimistic that this architecture can eventually run on newer Macs too.
There Are Still Lots of Problems
This is very far from a finished world model. The biggest obvious problem is still long-term consistency.. Objects can change a lot. The geometry of the environment can drift. Sometimes the model misunderstands what part of the scene should remain fixed. Even in the above demo, you can see it occasionally get confused about the environment.
In the next iteration, I'm hoping to change all that.
What's Next
My goal is to keep improving:
- long-term consistency
- text controllability
- action responsiveness
- image quality
- longer context
- lower inference cost
- Mac / MLX inference
- lower-end RTX performance
And eventually I want to release something where you can give it an image, type something, press WASD, and just start exploring the generated world on your own computer.
I specifically used a similar kind of starting image to some of my earlier demos because I wanted to see how much the system had changed over the last few months.
There is still a lot of work left.
🤗✨
