Robot Learning Benchmarks
The benchmarks used to evaluate manipulation and policy-learning methods, explained.
CALVIN
CALVIN: Composing Actions from Language and Vision. Long-horizon, language-conditioned manipulation benchmark. RoboFlamingo. RCSV.
COLOSSEUM
COLOSSEUM: Large-scale real-robot manipulation benchmark. Diverse tasks. BridgeVLA 64%. RCSV.
Google Robot Benchmark
Google Robot Benchmark: 700+ real-world manipulation tasks. WidowX, multi-embodiment. Success rate evaluation. RCSV.
LIBERO
LIBERO: Lifelong learning benchmark. 130 tasks, RoboSuite. InternVLA 95.9%. RCSV.
RLBench
RLBench: 100+ manipulation tasks in PyRep simulation. VLA evaluation benchmark. BridgeVLA 88.2%, InternVLA. RCSV.







