0° tilt — succeeds
Real-to-sim evaluation for robot policies
Evaluate robot policies
before the lab.
Tracefield turns real scenes into simulation-backed evaluation suites, so teams can rank checkpoints, map failure boundaries, and compare generalist policies before running expensive real-world rollouts.
2-5 min
scene scan
1000s
parallel rollouts
r = 0.93
real-sim signal
Tracefield rollout
Close drawer · baseline
Simulated policy rollout under calibrated camera pose.
Every policy deploy today costs real robot hours. Teams ship checkpoints based on lab success rates — then discover lighting shifts, camera bumps, and placement variance on hardware. Tracefield evaluates policies offline, maps the exact operating envelope, and ranks checkpoints before a single real-robot hour is spent.
Interactive eval demo
See how Tracefield ranks policies, maps failure boundaries, and surfaces rollout evidence.
Explore a worked RT-1-X evaluation below — the same pipeline Tracefield runs on your scenes, checkpoints, and perturbation sweeps before hardware time is scheduled.
Parallel evaluation
Eight lighting presets · same checkpoint · same task
RT-1-X · close drawer · lighting sweep
Tracefield eval suite
Evaluate robot policies before the lab.
Tracefield reconstructs your scene, runs thousands of parallel rollouts, and returns rankings, robustness maps, and failure clips your team can act on. This demo walks through the RT-1-X worked example.
Watch controlled rollouts where behavior breaks — not just the headline score.
r = 0.93
Real-sim correlation
9 tasks
10°
Reliable camera tilt
close drawer
1000s
Parallel rollouts
per eval run
Camera tilt
Reliable through 10° on close-drawer, then the policy misses the handle.
15° tilt — fails
Lighting
Dim scenes collapse perception; brighter lighting recovers the task.
1500 lux (bright) — succeeds
100 lux (very dim) — fails
Everything you need to compare policies.
Nothing to rediscover on hardware.
Real scenes become eval environments.
Start with a short scan of the bench, kitchen, lab, or warehouse cell. Tracefield reconstructs geometry and appearance, inserts the robot, and turns the scene into a repeatable simulation suite.
scan
mesh
robot
eval
Rankings that track reality.
Simulated scores tracked published real-world task success in the worked example.
Find the cliff edge.
Sweep lighting, camera pose, object placement, and scene setup. The output is not one success number. It is the envelope where the policy still works.
Thousands of rollouts in parallel.
Evaluate many checkpoints, tasks, and perturbation sweeps without queueing the robot for every hypothesis. A full robustness map completes in minutes, not weeks.
1,024
rollouts per sweep
4
perturbation axes
12 min
full eval runtime
8 lighting presets in parallel
6 pass · 2 fail
Failure clips your team can act on.
Tracefield surfaces the rollout where behavior changes, not just the aggregate score. Review the failure clip, fix the policy, rerun the suite.
Close drawer · 0° camera tilt
Close drawer · 15° camera tilt
Same policy, same task — only the controlled perturbation changes. Tracefield flags the exact rollout where behavior breaks.
Built from reality
Simulation that starts
with your scene.
Generic sim benchmarks drift away from the real world. Tracefield uses reconstructed environments, task-specific objects, and controlled perturbations to make simulation a stronger signal for policy improvement.
scan -> reconstruction
reconstruction -> rollout suite
rollouts -> policy ranking
failures-> next data targetEarly access
Bring your next robot eval into simulation first.
Send us a policy, task, or environment you already evaluate on hardware. We'll show what Tracefield can measure before your next real-world rollout.