Practical Model Selection
Compare architectures by task fit, data need, and deployment complexity.
Side-by-side comparison of 50+ open-source vision-language-action (VLA) and vision-language (VLM) models for robotics — OpenVLA, Octo, RT-X, Diffusion Policy, LeRobot, and more — with benchmarks, licenses, hardware requirements, and paper links.
These pages capture users searching by deployment question, workflow, or commercial decision instead of a specific model name.
Foundation action models, trade-offs, and fit.
Workflow GuideWhat works best when demonstrations are your starting point.
Manipulation GuideForce, tactile signals, and recovery-aware policy choices.
Decision GuideBroad capability versus faster narrow deployment.
Decision GuideData, task scope, evaluation, and deployment constraints.
OpenArm GuidePolicy choices and practical starting paths for OpenArm.
Our team has deployed OpenVLA, Octo, and RT-X on OpenArm, Unitree G1, and Mobile ALOHA rigs. Tell us your robot and task and we will recommend a model and a dataset.
Get a Model Recommendation8 models match current filters.

7B-parameter VLA. Llama 2 + DINOv2/SigLIP. 970K demos from Open X-Embodiment. Outperforms RT-2-X with 7× fewer params. MIT, Hugging Face.
View model →
Transformer diffusion policy. 27M/93M params. 800K trajectories. Multi-robot, language/goal conditioning. MIT, Hugging Face.
View model →
Open X-Embodiment models. JAX & TensorFlow checkpoints. Multi-robot, language-conditioned. Apache 2.0.
View model →
Spatially guided VLA. Two-stage: grounding + action. 71–81% on Google Robot, 95.9% LIBERO. MIT, Hugging Face.
View model →
OpenFlamingo-based VLM for robot control. Policy head + imitation learning. Strong on CALVIN. MIT, Hugging Face.
View model →
3D VLA with input-output alignment. 88.2% RLBench, 64% COLOSSEUM. Heatmap pre-training + point cloud fine-tuning.
View model →
Visuomotor policy as denoising diffusion. +46.9% over prior methods. Receding horizon, time-series transformer. Open source.
View model →
Framework + ACT, SmolVLA (450M). End-to-end IL/RL. Datasets, training, deployment. PyTorch, Hugging Face Hub.
View model →Suggestions based on the model category you are exploring.
Compare architectures by task fit, data need, and deployment complexity.
Model choices are connected to compatible dataset and format stacks.
Open-source links and implementation-ready pointers reduce setup friction.
From evaluation to deployment with support for tuning and integration.
We provide data collection, fine-tuning support, and deployment for robot learning.
Every VLA here can be fine-tuned on your own demonstrations.