Comparable Metrics
Benchmarks are grouped for apples-to-apples performance checks.
Standardized evaluation for robot manipulation — RLBench, LIBERO, CALVIN, and more. Success rates, task completion, evaluation metrics.
Real-time rankings from Papers with Code, updated daily across manipulation, locomotion, and navigation tasks.
Loading leaderboard data…
5 benchmarks match current filters.
100+ manipulation tasks in PyRep. Widely used for VLA evaluation. BridgeVLA 88.2%, InternVLA 95%+ on subsets.
View benchmark →SimulationLifelong learning benchmark. 130 tasks, spatial/object/goal suites. RoboSuite. 95.9% SOTA (InternVLA).
View benchmark →SimulationComposing Actions from Language and Vision. Long-horizon, language-conditioned. RoboFlamingo strong baseline.
View benchmark →Real RobotReal-world manipulation. 700+ tasks. WidowX, various embodiments. Success rate, multi-task evaluation.
View benchmark →Real RobotLarge-scale real-robot benchmark. Diverse tasks, environments. BridgeVLA 64%.
View benchmark →Suggested stack based on selected benchmark type.
Benchmarks are grouped for apples-to-apples performance checks.
Evaluate both controlled and deployment-oriented settings.
Each benchmark path links to compatible model families.
Support for data capture and evaluation operations when needed.
We provide data collection and real-world evaluation support.
Benchmarks compare policies in sim; our eval loop scores them on physical rigs.