Robot learning
Four tests for a robot placement
A placement can reach the goal and fall over afterward. Four historical Samsara protocols tested when success should count.
A robot can place an object at the goal and knock it away on its next move. If the evaluator stops at the first success, both a stable placement and a brief contact count as passes.
Our earlier Samsara experiments used four stopping rules. The first counted success as soon as the environment reported the goal state. The second required it to persist while the policy kept acting. The third then stopped control and waited twenty steps. The fourth added actuator disturbance during the rollout, while retaining the sustained and hands-off requirements.
Each result covered 100 episodes per suite, for 400 total, with seed 1701 and initial states 10–19. The 96.50% first-success result used v9; 93.75% under sustained control used v10. The hands-off and disturbance results, 92.25% and 90.25%, used v12. The changed checkpoints prevent us from attributing the score differences solely to the grading rules.
The current benchmark uses a different rule
The current Samsara evaluation uses a 217M-parameter architecture and 2,000 episodes. Its pinned evaluation rollout stops at the first successful environment termination. The 94.85% overall score and 99.6% Object score use that rule.
A controlled comparison of the four earlier rules would hold the policy fixed and replay the same conditions. Until then, the historical results and the current benchmark answer different questions about manipulation.