Robot learning
Introducing Samsara-VLA
Robot control with memory, in two compact models.
From instruction to completion
Samsara-VLA above. Samsara-VLA Tiny below. Every category stays in view.
Spatial
Find the object described by its position.
Samsara-VLA
“Pick up the black bowl between the plate and the ramekin and place it on the plate”
Samsara-VLA Tiny
“Pick up the black bowl between the plate and the ramekin and place it on the plate”
Object
Select the named object among distractors.
Samsara-VLA
“Pick up the cream cheese and place it in the basket”
Samsara-VLA Tiny
“Pick up the cream cheese and place it in the basket”
Goal
Change the scene to match the instruction.
Samsara-VLA
“Open the middle drawer of the cabinet”
Samsara-VLA Tiny
“Open the middle drawer of the cabinet”
Long
Carry out an instruction with several actions.
Samsara-VLA
“Put both the alphabet soup and the cream cheese box in the basket”
Samsara-VLA Tiny
“Put both the alphabet soup and the cream cheese box in the basket”
Samsara-VLA is a robot policy that carries context from one observation to the next. It uses two camera views and an instruction to pick, place, open and complete sequences of actions. The videos above show complete successful episodes from the two released models.
Two sizes. One policy interface.
The 217M-parameter Samsara-VLA completed 1,890 of 2,000 native LIBERO episodes. Samsara-VLA Tiny uses 136M parameters and completed 1,786 under the same evaluation protocol.
Full model
Samsara-VLA
- Native LIBERO success
- 94.50%
- Parameters
- 217.43M
1,890 / 2,000 successful episodes
Compact model
Samsara-VLA Tiny
- Native LIBERO success
- 89.30%
- Parameters
- 136.52M
1,786 / 2,000 successful episodes
The difference is clearest on longer tasks: 89.8% success for the full model and 76.0% for Tiny. Object selection remains strong in both, at 99.4% and 98.2%. Tiny reduces the parameter count by 37.2%.
Context between actions
Each decision combines a fresh view of the scene with a compact memory of previous observations. The policy predicts twelve control actions, executes them, then observes again. Its memory resets when a new task begins.
A policy you can run locally
Both models share a Python SDK for CPU and CUDA execution, with ONNX exports for browser inference. The browser preview runs on your device using WebGPU or WebAssembly CPU after loading the model.
These are experimental simulation policies. The native results cover forty trained task templates with one evaluation seed; browser demonstrations are separate. Models and the preview are available through research access. Read the full model card for methods and limitations, or explore the research page for suite results and comparisons with published models.
MODEL DETAILS
Samsara-VLA
Robot control with recurrent memory. 94.50% native LIBERO success in a 217M-parameter policy.