奖励塑形
修改奖励函数以提供更密集、更信息丰富的学习信号,而不改变最优控制策略。例如,添加基于势函数的塑形奖励,为接近目标提供部分奖励。奖励塑形在稀疏奖励环境中能显著加速强化学习训练,但需要谨慎以避免引入意外的局部最优。

Open 7-DoF arm, bimanual-ready. Ships from San Francisco with pilot data-collection packages from $2,500.

90+ platforms — humanoids, arms, quadrupeds, dexterous hands, and teleoperation kits. In-stock hardware ships in 48 hours.

Hardware-synced demonstrations for VLA and imitation learning. Download in LeRobot, HDF5, or RLDS format.

From first arm on the bench to a policy that survives a real workcell — the Robotics Center pipeline in three steps.

Low-latency data-collection glove for dexterous manipulation. From $5,500, ships from San Francisco.

Unbox, calibrate, and run your first autonomous walk. Includes ROS 2 bringup and teleop quickstart.
修改奖励函数以提供更密集、更信息丰富的学习信号,而不改变最优控制策略。例如,添加基于势函数的塑形奖励,为接近目标提供部分奖励。奖励塑形在稀疏奖励环境中能显著加速强化学习训练,但需要谨慎以避免引入意外的局部最优。
