EN
Contact
Menu

Research Experimental

Samsara-VLA

Robot control with memory.

Two compact robot policies for language-conditioned manipulation. Watch them pick, place, open and complete multi-step tasks.

Samsara-VLA / In action
Two models. Four kinds of manipulation. Selected successful episodes from both released checkpoints. Scene and wrist cameras are shown together. Complete recorded actions, replayed in the original simulator. Playback follows simulation time and excludes inference waits. The full evaluation determines the success rates.

Two models, one interface

Samsara-VLA turns camera images, an English instruction and robot state into continuous control. It carries a compact memory between observations and predicts twelve actions at a time. Samsara-VLA Tiny uses the same interface with a smaller visual encoder.

Full model

Samsara-VLA

Native LIBERO success
94.50%
Parameters
217.43M

1,890 / 2,000 successful episodes

Compact model

Samsara-VLA Tiny

Native LIBERO success
89.30%
Parameters
136.52M

1,786 / 2,000 successful episodes

Tiny uses 37.2% fewer parameters, with a 5.2-point difference in overall success. Both counts include the frozen text encoder. Both models trained for 40,000 updates and 5.12 million supervised sample presentations.

Native LIBERO results

We evaluated each released checkpoint on 40 tasks and 2,000 episodes. Every task uses fifty official initial states. Object selection is the strongest suite for both models; longer sequences show the largest difference between them.

Samsara-VLASamsara-VLA Tiny

Spatial

Samsara-VLA 94.2%471 / 500
Samsara-VLA Tiny 89.8%449 / 500

Object

Samsara-VLA 99.4%497 / 500
Samsara-VLA Tiny 98.2%491 / 500

Goal

Samsara-VLA 94.6%473 / 500
Samsara-VLA Tiny 93.2%466 / 500

Long

Samsara-VLA 89.8%449 / 500
Samsara-VLA Tiny 76.0%380 / 500
Success rate by task suite. Every bar runs from 0 to 100%; counts show completed episodes out of 500 trials.

The full model completes 449 of 500 Long episodes; Tiny completes 380. On Object, they complete 497 and 491 respectively. These counts make the size–capability tradeoff more useful than a single overall score.

This is a self-reported, single-seed simulation evaluation on trained task templates, not an independent holdout. Physical robots and unseen tasks have not been validated.

Model size and capability

A compact policy makes local deployment easier to explore. The table places our two released models alongside published LIBERO results, with model size shown for every row. It does not establish a latency or memory-use ranking.

Model size and LIBERO success
ModelParametersOverall success
Samsara-VLAOne policy · 50 trials per task · our evaluation217.43M94.50%
Samsara-VLA TinyOne policy · 50 trials per task · our evaluation136.52M89.30%
OpenVLA-OFT · unified 2 views + state · one policy · published7B base96.8%
π₀ 2 views + state · published3.3B94.2%
SmolVLA · 450M Multi-task · 10 trials/task · replan each action450M87.3%
OpenVLA 1 view · published7B76.5%

Published references, not matched reruns. Camera inputs, training data, policy sharing and trial counts differ. Parameter counts include frozen components for Samsara; reference sizes follow the cited papers.

How the policy acts

The instruction identifies what to do; the camera views show the current scene. Memory carries context into the next decision. Fresh visual features and that history jointly inform the next action sequence.

Samsara-VLA / Observe, remember, act

Current vision, persistent history.

  1. 01 / Observe

    Read the scene from two views.

    The scene camera locates the cabinet. The wrist camera shows the view from the gripper. Both 256 × 256 images and the instruction enter the policy at each replan.

    “Open the middle drawer of the cabinet.”

    Scene cameraWrist camera
    Scene and wrist views at the start of the recorded drawer task.
    Camera views + instruction + robot state

    ↻ Observe again. Keep history until the episode ends.

    Scene and wrist views from the same recorded episode.
  2. 02 / Remember

    Update history without losing spatial detail.

    Two recurrent blocks summarize earlier observations and robot state. One history token reaches the action decoder alongside the current image features, which retain their spatial detail.

    “Open the middle drawer of the cabinet.”

    Earlier observations

    Compact recurrent state

    Current camera views + memory

    Context for the next decision

    ↻ Observe again. Keep history until the episode ends.

    A schematic of retained context, not a visualization of learned memory values.
  3. 03 / Act

    Predict a short sequence of actions.

    The decoder predicts twelve actions, each with seven control values for position, rotation and the gripper. The controller executes the chunk before asking the policy for another one.

    “Open the middle drawer of the cabinet.”

    One prediction / Twelve actions

    01
    02
    03
    04
    05
    06
    07
    08
    09
    10
    11
    12
    PositionRotationGripper

    Action structure shown schematically. Each column is one control step.

    ↻ Observe again. Keep history until the episode ends.

    Twelve actions, each with seven control values. The diagram shows the output structure.
  4. 04 / Repeat

    Observe again, with history intact.

    New camera images update the current view. The history cache keeps the same shape as the episode grows and resets at the next episode. Its effect on success still needs a matched ablation.

    “Open the middle drawer of the cabinet.”

    Scene cameraWrist camera
    Scene and wrist views after the model completes the drawer task.
    Fresh observations + retained history

    ↻ Observe again. Keep history until the episode ends.

    The drawer task completes in this recorded example. The policy resets its history for the next episode.

Camera frames come from the released model’s drawer episode. Memory and action diagrams show the interface; they do not measure a causal benefit from memory.

Runs on your device

The same policy can run through the Python SDK on CPU or CUDA, or through ONNX in a browser. Browser execution uses WebGPU when available and WebAssembly CPU otherwise. After the model loads, inference runs locally without a prediction service.

The browser preview is a separate demonstration from the native benchmark. In fixed-scene CPU checks, the full model completed both the bowl and drawer tasks; Tiny completed the bowl task and reached the time limit on the drawer. These checks are not browser success-rate or latency measurements.

Models and the browser preview currently require research access. EUPE-derived weights retain the FAIR Noncommercial Research License. The full model card covers the protocol, architecture and deployment requirements.

Research access

Work with Samsara.

Contact us for model access, the browser preview or a research collaboration.

research@goodailabs.com