稀疏奖励
仅在任务完成时(成功/失败)提供信号的奖励函数,中间步骤的奖励为零。稀疏奖励是最自然的规范(二元成功),但对强化学习来说最难学习——智能体必须通过随机探索发现成功行为。HER、课程学习和奖励塑形可以解决稀疏奖励学习问题。

Open 7-DoF arm, bimanual-ready. Ships from San Francisco with pilot data-collection packages from $2,500.

90+ platforms — humanoids, arms, quadrupeds, dexterous hands, and teleoperation kits. In-stock hardware ships in 48 hours.

Hardware-synced demonstrations for VLA and imitation learning. Download in LeRobot, HDF5, or RLDS format.

From first arm on the bench to a policy that survives a real workcell — the Robotics Center pipeline in three steps.

Low-latency data-collection glove for dexterous manipulation. From $5,500, ships from San Francisco.

Unbox, calibrate, and run your first autonomous walk. Includes ROS 2 bringup and teleop quickstart.
仅在任务完成时(成功/失败)提供信号的奖励函数,中间步骤的奖励为零。稀疏奖励是最自然的规范(二元成功),但对强化学习来说最难学习——智能体必须通过随机探索发现成功行为。HER、课程学习和奖励塑形可以解决稀疏奖励学习问题。
