EVERWORLD

The next scene,shaped by you

Real-time interactive spatial video models

Explore a scene, act on its objects and talk to its characters. Your input shapes what happens next.

From a workshop to connected harbor spaces and autonomous characters: 1.0 Play, 2.0 Compose, 3.0 Evolve.

INTERACTIVE ENTERTAINMENT

The freedom to explore is yours

Interactive livestreams

Audiences choose together, turning the story in a new direction.

Playable stories

Enter a role. Actions and dialogue change the next scene.

Versions

Play · Compose · Evolve

1.0

Play

Participant

Interactive scene generation

Explore a room-scale scene with your audience: ring the bell, unfold the map and get a response. Audiovisual generation continues, and your changes persist when you look away and return.

2.0

Compose

Creator

Large-scale spatial orchestration

Describe a larger scene, then shape connected streets, docks and a lighthouse. Create interactive objects, edit the layout and rules, and step inside to try the space you designed.

3.0

Evolve

Collaborator

Agent-driven scene evolution

Talk to your guide while exploring and change the route mid-conversation. Characters remember their experience, weigh possible actions and pursue their goals together as the scene evolves.

Architecture

Product capabilities and shared foundation

1.0 interactive generation, 2.0 spatial orchestration and 3.0 agent-driven evolution share persistent state, spatial perception and joint audiovisual generation. Candidate actions are evaluated in isolated branches before live execution.

1.0 / Interactive generation

Continuous, coherent interaction

Interact within a bounded scene. Keep state, geometry, sound and visuals consistent as the scene changes.

Bounded-space interaction · State and geometry consistency · Continuous AV

2.0 / Spatial orchestration

Create, edit and execute scenes

Build connected spaces, edit locally and run physics, skills and events through explicit scene rules.

Multimodal creation · Executable rules · Verification and revision

3.0 / Agent-driven evolution

Characters with memory, goals and decisions

Talk to characters in real time. Their memory, goals and observations guide the actions they choose.

Live multimodal interaction · Role memory and goals · Action planning

One shared foundation

Persistent state and memory, spatial perception and geometry, multimodal understanding and joint AV generation. Established in 1.0 and extended by 2.0 and 3.0.

Plan first, then act

Candidate actions → isolated rollouts → evaluate and select → live execution. 3.0 reuses the 2.0 executor; planning does not modify the live scene.

Product capabilities and shared foundationOpen full diagram ↗
Read the technical guide

Join us

Build the next chapter with us

Contribute to the models, runtime, evaluation or interactive experiences.

wenxiang@zju.edu.cn

Sound is muted by default. Use the player to enable it.

Production notes

Scenes and sound: GPT Image 2 and Seedance. Interface animation, dialogue and subtitles are composed for the product walkthrough.