RL² Eval

Prove it on real hardware

Close the loop from hardware to data to evaluation. RL² Eval benchmarks robot policies on real platforms so you can see what actually works before you deploy.

RL² is our hardware → data → evaluation stack. Eval is the last leg: run candidate policies on the same platforms we sell and the datasets we collect, then compare success rates, robustness, and failure modes on real robots instead of simulation alone. Evaluation slots are scheduled with our engineering team — tell us the embodiment and task and we’ll scope a run.

What you get

Results you can act on — not a leaderboard number

Every run produces structured, replayable evidence of how a policy behaves on the platforms we operate in-house.

Real-hardware runs

Policies execute on the same platforms we sell and the datasets we collect — not simulation alone.

Failure taxonomy

Success rate plus a breakdown of where and why a policy fails, so you know what to fix.

Replayable episodes

Every scenario is captured so you can rewatch, compare runs, and share evidence with your team.

How it works

RL² is our hardware → data → evaluation stack. Eval is the last leg.

  1. Pick a platform & task

    Choose an embodiment and task from the hardware we operate in-house.

  2. Run against held-out scenarios

    Your policy runs against hardware-synced demonstrations and unseen scenarios.

  3. Get structured results

    Receive success rates, a failure taxonomy, and replayable episodes you can compare over time.