# Samsara-VLA · Good AI Labs > Two compact robot policies for language-conditioned manipulation. Watch them pick, place, open and complete multi-step tasks. URL: https://www.goodailabs.com/research/samsara-vla/ Date: 2026-10-05 Model card: https://huggingface.co/gai-labs/samsara-vla/blob/main/MODEL_CARD.md Two models. Four kinds of manipulation. Selected successful episodes from both released checkpoints. Scene and wrist cameras are shown together. Complete recorded actions, replayed in the original simulator. Playback follows simulation time and excludes inference waits. The full evaluation determines the success rates. Spatial: Find the object described by its position. Samsara-VLA · Spatial. Complete recorded evaluation episode from the released checkpoint. Selected successful example. Playback follows simulation time and excludes inference waits. Recording: https://www.goodailabs.com/media/samsara/release/base/spatial.mp4 Samsara-VLA Tiny · Spatial. Complete recorded evaluation episode from the released checkpoint. Selected successful example. Playback follows simulation time and excludes inference waits. Recording: https://www.goodailabs.com/media/samsara/release/tiny/spatial.mp4 Object: Select the named object among distractors. Samsara-VLA · Object. Complete recorded evaluation episode from the released checkpoint. Selected successful example. Playback follows simulation time and excludes inference waits. Recording: https://www.goodailabs.com/media/samsara/release/base/object.mp4 Samsara-VLA Tiny · Object. Complete recorded evaluation episode from the released checkpoint. Selected successful example. Playback follows simulation time and excludes inference waits. Recording: https://www.goodailabs.com/media/samsara/release/tiny/object.mp4 Goal: Change the scene to match the instruction. Samsara-VLA · Goal. Complete recorded evaluation episode from the released checkpoint. Selected successful example. Playback follows simulation time and excludes inference waits. Recording: https://www.goodailabs.com/media/samsara/release/base/goal.mp4 Samsara-VLA Tiny · Goal. Complete recorded evaluation episode from the released checkpoint. Selected successful example. Playback follows simulation time and excludes inference waits. Recording: https://www.goodailabs.com/media/samsara/release/tiny/goal.mp4 Long: Carry out an instruction with several actions. Samsara-VLA · Long. Complete recorded evaluation episode from the released checkpoint. Selected successful example. Playback follows simulation time and excludes inference waits. Recording: https://www.goodailabs.com/media/samsara/release/base/long.mp4 Samsara-VLA Tiny · Long. Complete recorded evaluation episode from the released checkpoint. Selected successful example. Playback follows simulation time and excludes inference waits. Recording: https://www.goodailabs.com/media/samsara/release/tiny/long.mp4 ## Two models, one interface Samsara-VLA turns camera images, an English instruction and robot state into continuous control. It carries a compact memory between observations and predicts twelve actions at a time. Samsara-VLA Tiny uses the same interface with a smaller visual encoder. Samsara-VLA: 2.17428023e+08 parameters; 94.50% overall success (1890/2000). - libero_10: 89.8% (449/500). - libero_goal: 94.6% (473/500). - libero_object: 99.4% (497/500). - libero_spatial: 94.2% (471/500). Samsara-VLA Tiny: 1.36519607e+08 parameters; 89.30% overall success (1786/2000). - libero_10: 76.0% (380/500). - libero_goal: 93.2% (466/500). - libero_object: 98.2% (491/500). - libero_spatial: 89.8% (449/500). Tiny uses **37.2% fewer parameters**, with a **5.2-point difference** in overall success. Both counts include the frozen text encoder. Both models trained for 40,000 updates and 5.12 million supervised sample presentations. ## Native LIBERO results We evaluated each released checkpoint on **40 tasks and 2,000 episodes**. Every task uses fifty official initial states. Object selection is the strongest suite for both models; longer sequences show the largest difference between them. Samsara-VLA: 2.17428023e+08 parameters; 94.50% overall success (1890/2000). - libero_10: 89.8% (449/500). - libero_goal: 94.6% (473/500). - libero_object: 99.4% (497/500). - libero_spatial: 94.2% (471/500). Samsara-VLA Tiny: 1.36519607e+08 parameters; 89.30% overall success (1786/2000). - libero_10: 76.0% (380/500). - libero_goal: 93.2% (466/500). - libero_object: 98.2% (491/500). - libero_spatial: 89.8% (449/500). Samsara-VLA native LIBERO: Spatial 94.2% (471/500); Object 99.4% (497/500); Goal 94.6% (473/500); Long 89.8% (449/500). Overall 94.50% (1,890/2,000). Release evidence: /research/samsara-vla/release-evaluation.json The full model completes 449 of 500 Long episodes; Tiny completes 380. On Object, they complete 497 and 491 respectively. These counts make the size–capability tradeoff more useful than a single overall score. This is a self-reported, single-seed simulation evaluation on trained task templates, not an independent holdout. Physical robots and unseen tasks have not been validated. ## Model size and capability A compact policy makes local deployment easier to explore. The table places our two released models alongside published LIBERO results, with model size shown for every row. It does not establish a latency or memory-use ranking. Model size and published LIBERO success: - OpenVLA-OFT · unified: 7B base parameters, 96.8% overall success. 2 views + state · one policy · published Source: https://arxiv.org/html/2502.19645v2 - π₀: 3.3B parameters, 94.2% overall success. 2 views + state · published Source: https://arxiv.org/html/2502.19645v2#S5.T1 - SmolVLA · 450M: 450M parameters, 87.3% overall success. Multi-task · 10 trials/task · replan each action Source: https://arxiv.org/html/2506.01844v1 - OpenVLA: 7B parameters, 76.5% overall success. 1 view · published Source: https://arxiv.org/html/2406.09246v3#A5.T12 Published references, not matched reruns. Camera inputs, training data, policy sharing and trial counts differ. Parameter counts include frozen components for Samsara; reference sizes follow the cited papers. ## How the policy acts The instruction identifies what to do; the camera views show the current scene. Memory carries context into the next decision. Fresh visual features and that history jointly inform the next action sequence. Current vision, persistent history. Camera frames come from the released model’s drawer episode. Memory and action diagrams show the interface; they do not measure a causal benefit from memory. 1. Read the scene from two views. The scene camera locates the cabinet. The wrist camera shows the view from the gripper. Both 256 × 256 images and the instruction enter the policy at each replan. Figure: Scene and wrist views from the same recorded episode. 2. Update history without losing spatial detail. Two recurrent blocks summarize earlier observations and robot state. One history token reaches the action decoder alongside the current image features, which retain their spatial detail. Figure: A schematic of retained context, not a visualization of learned memory values. 3. Predict a short sequence of actions. The decoder predicts twelve actions, each with seven control values for position, rotation and the gripper. The controller executes the chunk before asking the policy for another one. Figure: Twelve actions, each with seven control values. The diagram shows the output structure. 4. Observe again, with history intact. New camera images update the current view. The history cache keeps the same shape as the episode grows and resets at the next episode. Its effect on success still needs a matched ablation. Figure: The drawer task completes in this recorded example. The policy resets its history for the next episode. ## Runs on your device The same policy can run through the Python SDK on CPU or CUDA, or through ONNX in a browser. Browser execution uses WebGPU when available and WebAssembly CPU otherwise. After the model loads, inference runs locally without a prediction service. The browser preview is a separate demonstration from the native benchmark. In fixed-scene CPU checks, the full model completed both the bowl and drawer tasks; Tiny completed the bowl task and reached the time limit on the drawer. These checks are not browser success-rate or latency measurements. Models and the browser preview currently require research access. EUPE-derived weights retain the FAIR Noncommercial Research License. The [full model card](https://huggingface.co/gai-labs/samsara-vla/blob/main/MODEL_CARD.md) covers the protocol, architecture and deployment requirements. Contact: research@goodailabs.com