返回Glossary
奖励工程
设计标量奖励函数的过程,该函数编码强化学习训练中所需的机器人行为。好的奖励工程提供密集反馈来指导学习,而不引入意外的捷径。常见陷阱:奖励黑客(优化奖励代理而非预期行为)和稀疏奖励(学习信号过于稀少)。

Open 7-DoF arm, bimanual-ready. Ships from San Francisco with pilot data-collection packages from $2,500.

90+ platforms — humanoids, arms, quadrupeds, dexterous hands, and teleoperation kits. In-stock hardware ships in 48 hours.

Hardware-synced demonstrations for VLA and imitation learning. Download in LeRobot, HDF5, or RLDS format.

From first arm on the bench to a policy that survives a real workcell — the Robotics Center pipeline in three steps.

Low-latency data-collection glove for dexterous manipulation. From $5,500, ships from San Francisco.

Unbox, calibrate, and run your first autonomous walk. Includes ROS 2 bringup and teleop quickstart.
设计标量奖励函数的过程,该函数编码强化学习训练中所需的机器人行为。好的奖励工程提供密集反馈来指导学习,而不引入意外的捷径。常见陷阱:奖励黑客(优化奖励代理而非预期行为)和稀疏奖励(学习信号过于稀少)。
