# Selecting Samsara's checkpoint ยท Good AI Labs > The selected 80,000-step checkpoint finished two episodes ahead of the 34,000-step candidate. All five results from the same 2,000-episode evaluation. URL: https://www.goodailabs.com/blog/a-number-you-cannot-tune-for/ Date: 2026-10-03 Samsara's selected 80,000-step checkpoint completed 1,897 LIBERO episodes. The 34,000-step checkpoint completed 1,895. Two successes separated them out of 2,000 attempts. We evaluated five candidates on the same forty tasks and fifty initial states per task, then selected the highest score. | Training updates | Successful episodes | Success rate | | ---: | ---: | ---: | | 20,000 | 1,862 / 2,000 | 93.10% | | 34,000 | 1,895 / 2,000 | 94.75% | | 40,000 | 1,878 / 2,000 | 93.90% | | 60,000 | 1,889 / 2,000 | 94.45% | | 80,000 | 1,897 / 2,000 | 94.85% | The scores did not improve steadily with training. A 0.10-point lead over the 34,000-step candidate is too small for this single-seed run to establish a reliable improvement. The [Samsara report](/research/samsara-vla/) includes the suite breakdown and training setup. Matching tasks and initial-state indices reveals more than the totals. The selected checkpoint succeeded on 68 starts that the 34,000-step checkpoint failed; the earlier checkpoint succeeded on 66 that the selected one failed. The two-episode lead contains 134 changed outcomes. ## Selection changes the question These episodes helped choose the model, so the 94.85% score is a development result. Confirmation would fix the checkpoint and evaluation procedure before using reserved examples, then repeat across additional seeds. The same issue applies to reflex-1's diagnostics. Its 95.0% Banking77 and 92.2% Emotion scores came from 500-example subsets inspected during development, on task families included in training. Those results describe familiar-task performance. A new application's questions need their own evaluation. Contact: research@goodailabs.com