Ground truth,
read from the engine.
4D capture from AAA game engines — pixels paired with the exact world state that produced them, tick by tick, plus human visual-reasoning chains over what happens next and why.
Three principles.
Engine-measured ground truth
Positions, velocities, contacts, camera poses and events are read directly from the engine at 60 Hz, tick-aligned to the rendered frames. Nothing is estimated from pixels after the fact — the supervision signal is the simulator's own state, which makes labels exact rather than approximate.
Counterfactual by construction
Because we control the engine, we can replay the same world and branch it — same scene, different action, divergent outcome. That turns "what happens next" from a single trajectory into a distribution a model can actually learn dynamics from, which passive video can never provide.
Diversity across worlds
Physics, lighting, scale and dynamics vary enormously across titles and genres — driving, flight, open-world, simulation. We capture across that spread deliberately, and pair it with human chain-of-thought visual reasoning, so models see both how worlds behave and how people reason about them.
How a capture is made.
Hooks at the engine level
Each title is instrumented so world state can be read as it's computed — entities, physics, camera, events — rather than reconstructed from the rendered output.
Tick-aligned, frame-exact
Video and state stream on the same clock at 60 Hz. Every frame carries the exact state that produced it, and branches are replayed from identical seeds for counterfactual pairs.
Human reasoning on top
Annotators write chain-of-thought visual reasoning over the captures — prediction, causality, intent — reviewed against the engine state, which makes the reasoning checkable.
Capture programs.
4D game capture
Gameplay video across AAA titles with full telemetry — real play, real dynamics, captured at scale across genres and physics regimes.
Engine-native world state
The state stream itself — entities, physics, camera and events at 60 Hz, tick-aligned to frames — the supervision layer for dynamics and prediction.
Visual-reasoning chains
Human chain-of-thought over captured scenes — prediction, causality, spatial reasoning — written by annotators and checked against engine ground truth.
Custom capture runs
Scenario, title mix, state schema and branching strategy scoped to your training objective — on the same instrumented capture pipeline.
From overview to catalog.
This page — capture programs, methodology, and what differentiates the data.
Representative captures with video, state streams and reasoning chains, on request.
A scoped capture run against your schema and objectives, evaluated before scale.
The full browsable world-model catalog — title inventories, scale and specs — customer access.
The full catalog — title inventories, scale and specs — is available to customers. Request access and we'll respond within one business day.
See the data behind this page.
We'll share representative captures, state-schema documentation, and scoping for custom capture runs.