AgentGartenCode Worlds for Evolving Agents

Executable environments with a real-time neural renderer

The idea

Code keeps the world.
The renderer shows it.

We keep each world as code: objects, rules and goals live in an executable program. A learned neural renderer turns the geometry that program exports into what the agent sees next, in real time. Drag across the frame to compare the two.

From geometry to observationDrag to compare
Untextured style1

The film

Step inside the worlds

Evolving agents

Agents that improve from experience

Pretrained agents play inside these worlds, seeing only rendered camera frames. After each round they write down what they learned, and the next round starts from those notes.

Build shelters

4rounds played
our visual agents

≈25Mtraining episodes
OpenAI self-play RL

Use ramps to enter shelters

10rounds played
our visual agents

≈100Mtraining episodes
OpenAI self-play RL

The overhead view of a game, with the agent's own first-person view in the corner.
Build coverA hider moves a panel to rebuild cover.
Find a wayA seeker carries a ramp over the wall.
Learn and retryA seeker moves the ramp closer and retries.

How a round works

Agent Loop

  1. 1Read

  2. 2Play

  3. 3Write

  4. 4Archive

Next round: new agents start from the task and every earlier playbook

The same loop in other worlds

Companion dog

Keep a dog willingly engaged for a 60-second session by offering a hand, petting, and playing with a ball.

Engagement score

13Round 114Round 219Round 319Round 4

By round 4 the agent alternates an offered hand and two chest strokes with brief ball play, pausing in between.

One-lane bridge

Two cars, each driven by its own agent from the windshield view, must swap ends of a bridge that fits one car.

Seconds until both cars arrive · lower is better

71 sRound 168 sRound 245 sRound 341 sRound 4

Blue learns to yield on the bank, turn while red passes, and re-center on the bridge without waiting.

Herding

Two dogs, seeing only from their own eye height, guide four sheep into a pen and hold them there for five seconds.

Score · out of 100

60Round 190.1Round 287.6Round 388.3Round 4

Round 1 timed out with three sheep penned. Rounds 2 to 4 penned all four without startling the flock.

Quarry loader

A wheel loader must push two rocks onto staging pads, deliver one to a bunker behind a wall, and park, within 360 seconds.

Score · out of 100

30Round 10Round 230Round 390.9Round 4

Round 1 cleared the rocks and got no further. Round 4 cleared them, delivered one and parked, with 31 seconds to spare.

World gallery

A playground of possibilities

How it works

From an action to the next frame

The interaction loop: the agent sends actions to the code world, where the engine updates the scene and exports depth or surface normals as conditions; the neural renderer turns them, together with a reference image, text and its visual history, into the next observation, which returns to the agent.
Real-time renderer · native speed

It runs in real time

The renderer generates video in short blocks, and each block is ready before the previous one has finished playing. An agent, or a person at the keyboard, sees the result of an action before choosing the next.

The renderer starts from a pretrained video model, learns to follow geometry, and is distilled into a block-by-block student that keeps up with the agent. Read how it is built in the blog

BibTeX

@misc{mirros2026evolvingagents,
  title  = {{AgentGarten}: Code Worlds for Evolving Agents},
  author = {{MirroS Team}},
  year   = {2026},
  month  = {Sep},
  url    = {https://mirros.ai/blog/worlds-for-evolving-agents},
  note   = {Blog post}
}