Open 7-DoF arm, bimanual-ready. Ships from San Francisco with pilot data-collection packages from $2,500.
90+ platforms — humanoids, arms, quadrupeds, dexterous hands, and teleoperation kits. In-stock hardware ships in 48 hours.
Hardware-synced demonstrations for VLA and imitation learning. Download in LeRobot, HDF5, or RLDS format.
From first arm on the bench to a policy that survives a real workcell — the Robotics Center pipeline in three steps.
Low-latency data-collection glove for dexterous manipulation. From $5,500, ships from San Francisco.
Unbox, calibrate, and run your first autonomous walk. Includes ROS 2 bringup and teleop quickstart.
信任域策略优化——一种策略梯度强化学习算法,通过信任域(由 KL 散度衡量)约束每次策略更新,以防止大幅破坏性更新。TRPO 提供单调改进保证,但由于约束优化计算成本高。PPO 用更简单的裁剪代理目标近似 TRPO 的优势。