策略蒸馏
通过训练学生策略匹配教师策略的动作分布(通过KL散度最小化),将大型或复杂的策略(教师)压缩为更小、更快的策略(学生)。策略蒸馏用于:压缩强化学习策略以实现实时部署、从仿真迁移到真实硬件,以及将多个专门化策略合并为一个。

Open 7-DoF arm, bimanual-ready. Ships from San Francisco with pilot data-collection packages from $2,500.

90+ platforms — humanoids, arms, quadrupeds, dexterous hands, and teleoperation kits. In-stock hardware ships in 48 hours.

Hardware-synced demonstrations for VLA and imitation learning. Download in LeRobot, HDF5, or RLDS format.

From first arm on the bench to a policy that survives a real workcell — the Robotics Center pipeline in three steps.

Low-latency data-collection glove for dexterous manipulation. From $5,500, ships from San Francisco.

Unbox, calibrate, and run your first autonomous walk. Includes ROS 2 bringup and teleop quickstart.
通过训练学生策略匹配教师策略的动作分布(通过KL散度最小化),将大型或复杂的策略(教师)压缩为更小、更快的策略(学生)。策略蒸馏用于:压缩强化学习策略以实现实时部署、从仿真迁移到真实硬件,以及将多个专门化策略合并为一个。
