Prove it on real hardware
Close the loop from hardware to data to evaluation. RL² Eval benchmarks robot policies on real platforms so you can see what actually works before you deploy.
RL² is our hardware → data → evaluation stack. Eval is the last leg: run candidate policies on the same platforms we sell and the datasets we collect, then compare success rates, robustness, and failure modes on real robots instead of simulation alone. Evaluation slots are scheduled with our engineering team — tell us the embodiment and task and we’ll scope a run.
Results you can act on — not a leaderboard number
Every run produces structured, replayable evidence of how a policy behaves on the platforms we operate in-house.
Real-hardware runs
Policies execute on the same platforms we sell and the datasets we collect — not simulation alone.
Failure taxonomy
Success rate plus a breakdown of where and why a policy fails, so you know what to fix.
Replayable episodes
Every scenario is captured so you can rewatch, compare runs, and share evidence with your team.
How it works
RL² is our hardware → data → evaluation stack. Eval is the last leg.
Pick a platform & task
Choose an embodiment and task from the hardware we operate in-house.
Run against held-out scenarios
Your policy runs against hardware-synced demonstrations and unseen scenarios.
Get structured results
Receive success rates, a failure taxonomy, and replayable episodes you can compare over time.







