Back to Benchmarks

CALVIN

CALVIN: Composing Actions from Language and Vision. Long-horizon, language-conditioned manipulation benchmark. RoboFlamingo. RCSV.

← Benchmarks

Composing Actions from Language and Vision — long-horizon, language-conditioned manipulation.

Overview

CALVIN evaluates language-conditioned manipulation over long horizons. Agents must compose multiple skills from natural language instructions. Simulation-based. RoboFlamingo and other VLM-based policies show strong performance.

Official Links

Benchmarks compare policies in sim; our eval loop scores them on physical rigs.

RCSV Eval → License training data →