Research Public weights
Reflex-1
Decision models for agents
A 421M-parameter decision model for agents. One forward pass, on CPU or GPU.
Fine-tuned weights available on Hugging Face.
Overview
Many agent decisions end with a choice: which queue receives a request, which tool fits a task, or which action comes next. Reflex-1 scores the choices your application supplies in a single forward pass. It returns a probability distribution that the application can use to make that decision.
The research follows the decision-model direction explored by Jev, with a distinct architecture and supervised training procedure. A 28-layer context encoder and a six-layer candidate encoder feed a shared scoring head. The current checkpoint updates all 420,778,370 parameters.
The latest fine-tuned checkpoint is the default public download. Its model card contains current evaluation results and CPU measurements. The dated studies and recordings below retain their original checkpoint identities.
Decision model
Public release · runtime 1.0.0- Use
- Choose among options supplied by an application: labels, routes or possible actions.
- Input → output
- State, question and 1–255 choices → a choice distribution in one forward pass.
- Model
- 421M parameters. A 28-layer context encoder, a 6-layer candidate encoder and a shared scoring head.
- Run locally
- CPU or GPU · FP32. The current CPU replay measured 2.17 GiB peak process memory with one or four compute threads.
- Evaluation
- 92.93% on a 495-case Banking77 reserved subset; 21/30 native MiniWoB++ episodes. Task families are represented in training.
- Access
- Public weights, tokenizers and inference code under Apache 2.0. English text; runtime 1.0.0.
The latest fine-tuned checkpoint is the default download. The dated studies and recordings below describe earlier checkpoints. The application supplies the choices and controls execution.
Current results
The October 5 release completed 21/30 native MiniWoB++ browser episodes with a fixed controller. It scored 92.93% on 495 Banking77 examples. These are reserved subsets from task families represented in training, not full benchmark scores.
| Task | Correct / evaluated | Accuracy |
|---|---|---|
| AG News | 465 / 500 | 93.00% |
| SST-5 | 299 / 500 | 59.80% |
| Emotion | 458 / 500 | 91.60% |
| Banking77 | 460 / 495 | 92.93% |
| BoolQ | 171 / 202 | 84.65% |
| MNLI | 436 / 500 | 87.20% |
| RACE | 287 / 500 | 57.40% |
| SciQ | 123 / 128 | 96.09% |
On an Intel Xeon Platinum 8558 CPU, Banking77 inference took 375.2 ms median and 452.1 ms p95 with four compute threads. Peak process memory was 2.17 GiB, including loading, warmup and inference.
Current CPU measurements
Intel Xeon Platinum 8558 · FP32 · batch size 1One compute thread
- AG News 1209.7 / 1385.0 ms
- SST-5 973.6 / 1160.3 ms
- Emotion 956.9 / 1165.8 ms
- Banking77 1101.7 / 1276.9 ms
- BoolQ 1600.8 / 2710.6 ms
Four compute threads
- AG News 389.5 / 480.4 ms
- SST-5 320.6 / 404.2 ms
- Emotion 318.8 / 391.0 ms
- Banking77 375.2 / 452.1 ms
- BoolQ 486.7 / 782.2 ms
480 calls per thread setting over 120 inputs in two fresh processes. Tokenization and inference included; loading and warmup excluded. Shared host, evaluation runtime, no cross-request caches.
The full evaluation is still incomplete. BoolQ, RACE and the public JevBench diagnostic regressed from the previous checkpoint; native Dino reached the 60-second target in 0/16 attempts. Current GPU latency has not been measured. The published evaluation records the checkpoint, conditions and remaining gaps.
How it works
Candidate answers are supplied with each request. The model supports 1–255 choices per question, within its token limits, and several questions can share a packed state. A support application can supply its own queues; an agent can supply descriptions of the tools it has available. New decision types need their own accuracy tests.
The model scores the supplied options; the application constructs tool arguments
and decides which actions are permitted. State text is capped at 512
context-tokenizer tokens and can be shortened further by the request budget.
Each choice accepts up to 512 option-tokenizer tokens. Inspect state_truncated
when context may have been shortened.
Reflex-1 / Decision model
From a question to a choice.
01 / Supply
Define the decision.
Your application supplies the state, a question and the choices it can act on. Those choices can change with every request: support queues, available tools or permitted actions.
State + question
The customer was charged twice.
Which issue matches?Context + questionModernBERTEach choiceMiniLMShared scoring headChoice probabilitiesApplicationPermission → arguments → actionIllustrative support request. The application defines the candidate set. 02 / Encode
Represent the context and the candidates.
An adapted ModernBERT encoder reads the state and question. A frozen MiniLM encoder represents each supplied choice. Several questions can share a packed state.
State + question
The customer was charged twice.
Which issue matches?Context + questionModernBERTEach choiceMiniLMShared scoring headChoice probabilitiesApplicationPermission → arguments → actionTwo encoders feed one shared scoring head. This is a schematic of the preview architecture. 03 / Score
Score every supplied choice in one pass.
The shared head combines context and candidate representations, then returns a probability distribution over the supplied choices. New decision types still need accuracy and calibration tests.
State + question
The customer was charged twice.
Which issue matches?Context + questionModernBERTEach choiceMiniLMShared scoring headChoice probabilitiesApplicationPermission → arguments → actionOne score per supplied candidate, normalized to a probability distribution. 04 / Act
Let the application carry out the decision.
The application decides which actions are permitted, constructs any tool arguments and executes the selected action. Reflex-1 supplies the scores; the surrounding system owns the control loop.
State + question
The customer was charged twice.
Which issue matches?Context + questionModernBERTEach choiceMiniLMShared scoring headChoice probabilitiesApplicationPermission → arguments → actionThe application boundary matters: scoring a tool does not execute it.
Architecture of the public development preview. The support request illustrates the input format.
Examples
The examples and decision-quality comparisons below use the October 3 snapshot. Speed and resource measurements are labelled separately. Each recording keeps its observed ending, including failure.
Chromium Dino
This recording uses Chromium’s original chrome://dino game. Reflex-1 receives
the currently visible obstacle geometry and chooses run, jump or duck. Native
keyboard events execute those choices while the game clock continues during
inference. The video shows actual browser frames at normal speed.
The two attempts with this snapshot ended in collisions after 42.31 seconds and 30.41 seconds; neither reached the 60-second limit. The longer attempt is shown first. These are independent random courses, so they do not support a paired model comparison. The complete second attempt remains available below.
Watch the second Dino attempt
ViZDoom
The game is native ViZDoom’s Defend the Center scenario at skill 5. Reflex-1 chooses actions from visible object bounds, health and ammunition. This is a structured-state controller; it does not infer actions from pixels. The recording below is the originally predeclared example, shown as a single Reflex-1 view.
Across all 32 reserved episodes, this snapshot averaged 18.88 kills and 22.47 seconds of survival. Every episode ended before the 45-second limit. The chart retains every evaluated external model and the full episode counts.
ViZDoom · kills and survival limits
3 October 2026 · development snapshot 70a843878214Defend the Center · skill 5
- Reflex-10/32 reached 45 s · 32 deaths · 0 errors · 22.47 s mean survival 18.995% CI 16.6–21.2
- Laya0/32 reached 45 s · 32 deaths · 0 errors · 4.90 s mean survival 1.495% CI 1.2–1.6
- OpenJev · DeBERTa0/32 reached 45 s · 32 deaths · 0 errors · 4.90 s mean survival 1.495% CI 1.2–1.6
- Kev-0.8B0/32 reached 45 s · 32 deaths · 0 errors · 4.90 s mean survival 1.495% CI 1.2–1.6
Reflex-1 denotes the October 3 development weights shown here, not the default download or a final release. Original ViZDoom engine, structured visible-state inputs and raw argmax actions; 45-second episode limit. Every declared outcome is retained. Mean intervals resample paired seeds 10,000 times. Reserved episode seeds do not establish unseen structured states. This named adapter/environment comparison is not an official benchmark score or a pixel-input policy evaluation.
Browser workflows
These are locally hosted synthetic WebGym sites distributed with Laya Browser. The browser states were captured in real Chromium, and the original task checker decides success. Reflex-1 selects actions and targets; a shared Qwen3-1.7B helper supplies typed text and dropdown values. The outcome measures that complete system.
The restaurant task succeeds; the shopping task fails. Both use the same predeclared task seed and retain the recorded ending. Playback is 4× the original wall time, replaying captured browser states with the terminal hold preserved. These examples are synthetic-site development results, not MiniWoB++, WebArena or live-site scores.
Browser completion · synthetic WebGym
3 October 2026 · development snapshot 70a843878214- Reflex-1485 decision calls (0 errors) · 102 helper calls (0 HTTP / 5 contract errors) 13/35
- Laya Browser505 decision calls (0 errors) · 100 helper calls (0 HTTP / 7 contract errors) 17/35
- Laya56 decision calls (0 errors) · 12 helper calls (0 HTTP / 0 contract errors) 0/35
- OpenJev · DeBERTa35 decision calls (24 errors) · 0 helper calls (0 HTTP / 0 contract errors) 0/35
- Kev-0.8B196 decision calls (0 errors) · 82 helper calls (0 HTTP / 0 contract errors) 1/35
Reflex-1 denotes the October 3 development weights shown here, not the default download or a final release. Real Chromium capture of local synthetic WebGym sites. Every arm uses the same executor, Qwen3-1.7B text helper and 40-step limit. Reflex-1 uses BF16 autocast and 1,024 state / 16,384 request tokens; other serving conditions are preserved in the download. All 35 attempts per arm, request failures and helper errors remain included. This panel is separate from MiniWoB++, WebArena and live-site evaluation.
Reflex-1 completed 13/35 tasks in this study; Laya Browser completed 17/35. All attempts, failures and model-request errors remain in the denominators. A later pilot through the official MiniWoB++ environment completed 0/30 tasks with this snapshot, so synthetic-site results should not be read as success on that benchmark.
Decision quality
Typed Decisions covers customer support, invoice processing, security alerts and agent traces. The evaluation preserves original state groups and asks each model the same five judgments through a common choice interface. Reflex-1 answered 1,508/2,000 judgments correctly (75.4%).
Structured decisions
3 October 2026 · development snapshot 70a843878214Customer support
- Reflex-1383/500 correct · 0 errors · 0 state cuts 76.6%95% CI 72.0–80.8%
- Laya · native210/500 correct · 0 errors · 39 state cuts 42.0%95% CI 38.2–45.8%
- OpenJev · DeBERTa206/500 correct · 0 errors · 350 state cuts 41.2%95% CI 37.8–44.0%
- Kev-0.8B314/500 correct · 0 errors · 10 state cuts 62.8%95% CI 58.2–67.4%
Invoice processing
- Reflex-1394/500 correct · 0 errors · 0 state cuts 78.8%95% CI 75.0–82.4%
- Laya · native192/500 correct · 0 errors · 0 state cuts 38.4%95% CI 32.6–44.6%
- OpenJev · DeBERTa203/500 correct · 0 errors · 500 state cuts 40.6%95% CI 38.2–43.0%
- Kev-0.8B252/500 correct · 0 errors · 0 state cuts 50.4%95% CI 45.2–55.8%
Security alerts
- Reflex-1362/500 correct · 0 errors · 0 state cuts 72.4%95% CI 68.4–76.2%
- Laya · native167/500 correct · 0 errors · 0 state cuts 33.4%95% CI 30.0–37.0%
- OpenJev · DeBERTa210/500 correct · 0 errors · 275 state cuts 42.0%95% CI 39.2–44.6%
- Kev-0.8B236/500 correct · 0 errors · 0 state cuts 47.2%95% CI 42.6–51.8%
Agent traces
- Reflex-1369/500 correct · 0 errors · 0 state cuts 73.8%95% CI 69.0–78.4%
- Laya · native193/500 correct · 0 errors · 0 state cuts 38.6%95% CI 33.4–44.0%
- OpenJev · DeBERTa192/500 correct · 0 errors · 0 state cuts 38.4%95% CI 34.6–42.2%
- Kev-0.8B179/500 correct · 0 errors · 0 state cuts 35.8%95% CI 31.6–40.4%
Reflex-1 denotes the October 3 development weights shown here, not the default download or a final release. 400 reserved source states, expanded to 2,000 judgments. Each workflow contains 500 judgments. Intervals resample source states, preserving the five related heads. Token limits: Reflex-1: state 1,024, request 4,096; Laya · native: request 512, instruction/choices 192; OpenJev · DeBERTa: state 256, request 512; Kev-0.8B: state 512, branch including state 2,048. Whole-cohort instruction/choice shortening (not additive across task charts): Laya · native 0/2,000. These counts are separate from the state cuts above.
On the reserved CLINC cohort, Reflex-1 answered 3,680/4,490 intent and out-of-scope requests correctly (82.0%). For unfamiliar-tool selection, it answered 72/87, below the strongest external result of 82/87. Both strengths and remaining gaps appear below.
Intent routing
3 October 2026 · development snapshot 70a843878214Known and unfamiliar requests
- Reflex-13,680/4,490 correct · 0 errors · 0 state cuts · 64 false rejections · 370 false acceptances 82.0%95% CI 80.8–83.1%
- Laya · native1,425/4,490 correct · 0 errors · 0 state cuts · 0 false rejections · 942 false acceptances 31.7%95% CI 30.3–33.2%
- Laya · expanded1,017/4,490 correct · 0 errors · 0 state cuts · 0 false rejections · 942 false acceptances 22.7%95% CI 21.4–23.9%
- OpenJev · DeBERTa2,496/4,490 correct · 0 errors · 0 state cuts · 5 false rejections · 939 false acceptances 55.6%95% CI 54.1–57.0%
- Kev-0.8B2,298/4,490 correct · 0 errors · 0 state cuts · 1 false rejections · 937 false acceptances 51.2%95% CI 49.6–52.6%
Reflex-1 denotes the October 3 development weights shown here, not the default download or a final release. 4,490 reserved requests: 3,548 known intents and 942 out-of-scope requests. Native Laya and its separately declared expanded-context control are both shown. Token limits: Reflex-1: state 1,024, request 4,096; Laya · native: request 512, instruction/choices 192; Laya · expanded: request 2,048, instruction/choices 1,536; OpenJev · DeBERTa: state 256, request 512; Kev-0.8B: state 512, branch including state 2,048. Whole-cohort instruction/choice shortening (not additive across task charts): Laya · native 4,490/4,490; Laya · expanded 0/4,490. These counts are separate from the state cuts above.
Selection among unfamiliar tools
3 October 2026 · development snapshot 70a843878214Tool selection
- Reflex-172/87 correct · 0 errors · 0 state cuts 82.8%95% CI 74.7–90.8%
- Laya · native80/87 correct · 0 errors · 14 state cuts 92.0%95% CI 86.2–97.7%
- Laya · expanded80/87 correct · 0 errors · 0 state cuts 92.0%95% CI 86.2–97.7%
- OpenJev · DeBERTa82/87 correct · 0 errors · 30 state cuts 94.3%95% CI 88.5–98.9%
- Kev-0.8B81/87 correct · 0 errors · 1 state cuts 93.1%95% CI 87.4–97.7%
Reflex-1 denotes the October 3 development weights shown here, not the default download or a final release. 87 reserved requests with tool names and normalized states separated before training. This is a tool-choice adaptation; argument generation and execution are outside its scope. Token limits: Reflex-1: state 1,024, request 4,096; Laya · native: request 512, instruction/choices 192; Laya · expanded: request 2,048, instruction/choices 1,536; OpenJev · DeBERTa: state 256, request 512; Kev-0.8B: state 512, branch including state 2,048. Whole-cohort instruction/choice shortening (not additive across task charts): Laya · native 87/87; Laya · expanded 65/87. These counts are separate from the state cuts above.
These are common-choice adaptations with declared cohorts, not official leaderboard submissions. External models were evaluated as shipped; they did not receive Reflex-1’s task-specific training. The figures retain the original uncertainty intervals, token budgets and truncation counts.
Speed and resources
The measurements in this section describe the October 2 checkpoint. Current checkpoint measurements are available in the model card linked above.
H200 request latency
We replayed 120 distinct requests twice per model on a shared NVIDIA H200. They came from five game and browser environments. Every model received the same inputs, with one request in flight at a time.
H200 request latency
120 inputs · two repeats · one request in flight- Reflex-1FP32 27.5 / 30.5
- OpenJevFP32 32.0 / 33.8
- Laya (native)BF16 27.1 / 142.0
- Laya (FP32 replay)FP32 40.2 / 44.4
Dot: median. Line end: p95. Includes tokenization, scoring and loopback HTTP; excludes loading and warmup.
Full measurement table and protocol
| Model | Compute | Median | p95 |
|---|---|---|---|
| reflex-1 | FP32 | 27.5 | 30.5 |
| OpenJev | FP32 | 32.0 | 33.8 |
| Laya (native) | BF16 | 27.1 | 142.0 |
| Laya (FP32 replay) | FP32 | 40.2 | 44.4 |
Laya native uses BF16 autocast; a separate FP32 replay is also included. OpenJev is the public DeBERTa checkpoint, not hosted Jev. These request timings do not measure task success or concurrent throughput. Recorded 2026-09-29.
reflex-1’s median was 27.5 ms and its p95 was 30.5 ms. It was faster than OpenJev in FP32 and Laya in the separate FP32 replay. Native Laya used BF16 autocast and had a slightly lower median. The shared machine and sample size limit what we can infer from close results.
The timer includes tokenization, model scoring, output normalization and loopback HTTP. It excludes loading and three warmup requests. Answer and option caches were disabled. We did not measure concurrent serving throughput in this run. OpenJev here is the public DeBERTa checkpoint; it is not the hosted Jev service.
M4 CPU latency
The CPU test used a MacBook Air M4, FP32, and batch size one. Each model used one PyTorch intra-op thread, one inter-op thread, and a BLAS thread limit of one. Each task had 24 distinct cases, each repeated twice. Model order was balanced across eight execution blocks.
M4 CPU request latency
FP32 · batch one · PyTorch threads 1/1 · BLAS limit 1Banking77
- Reflex-1 345.2 / 692.0
- OpenJev 1,332.6 / 2,857.1
- Laya 912.4 / 2,649.9
- Kev 4,283.4 / 11,073.7
Emotion
- Reflex-1 256.2 / 610.4
- OpenJev 381.8 / 1,029.5
- Laya 240.1 / 427.7
- Kev 1,121.9 / 4,422.3
AG News
- Reflex-1 321.7 / 578.7
- OpenJev 424.5 / 934.0
- Laya 324.1 / 760.5
- Kev 1,356.5 / 4,997.2
SST-5
- Reflex-1 299.1 / 631.0
- OpenJev 360.5 / 894.5
- Laya 250.7 / 548.5
- Kev 902.2 / 4,480.5
BoolQ
- Reflex-1 463.4 / 1,329.1
- OpenJev 451.9 / 1,398.9
- Laya 440.4 / 1,587.9
- Kev 1,775.3 / 6,548.0
Dot: median. Line end: p95. All tasks use the same logarithmic scale. Memory: 1.50 GiB peak process RSS in a separate Reflex-1 replay; not measured per model in this comparison.
Full measurement table and protocol
| Task | reflex-1 | OpenJev | Laya | Kev |
|---|---|---|---|---|
| Banking77 | 345.2/ 692.0 | 1332.6/ 2857.1 | 912.4/ 2649.9 | 4283.4/ 11073.7 |
| Emotion | 256.2/ 610.4 | 381.8/ 1029.5 | 240.1/ 427.7 | 1121.9/ 4422.3 |
| AG News | 321.7/ 578.7 | 424.5/ 934.0 | 324.1/ 760.5 | 1356.5/ 4997.2 |
| SST-5 | 299.1/ 631.0 | 360.5/ 894.5 | 250.7/ 548.5 | 902.2/ 4480.5 |
| BoolQ | 463.4/ 1329.1 | 451.9/ 1398.9 | 440.4/ 1587.9 | 1775.3/ 6548.0 |
Native runtime budgets. Laya truncates Banking77 options at its native budget; expanded-budget results are reported separately. Kev uses its CPU reference kernels and unmerged adapters. Recorded 2026-09-24.
Reflex-1 measured memory: 1.50 GiB peak process RSS. Separate Reflex-1 resource replay of the original 120 inputs, each scored twice. Whole-process peak RSS includes imports, model loading, warmup and inference. The original four-model latency comparison did not record per-model memory. Recorded 2026-10-02 UTC. Resource measurements and protocol.
Banking77 asks the model to choose among 77 intents. reflex-1 took 345.2 ms median; OpenJev took 1,332.6 ms in the same run. Across the five tasks, reflex-1’s medians ranged from 256.2 to 463.4 ms.
Laya was faster on several smaller-choice tasks. Its native Banking77 budget truncated the options, so its timing covers a reduced candidate set. Kev used CPU reference kernels with unmerged adapters. Background laptop activity contributed to the wider tail latencies.
This test includes request construction through normalized output, with the model already loaded. The CPU and GPU tests use different requests; their timings do not give a direct CPU-to-GPU speedup.
Familiar-task accuracy
reflex-1 scored 95.0% on Banking77 and 92.2% on Emotion. It led the compared checkpoints on both, tied the best score on AG News and SST-5, and scored slightly below OpenJev on BoolQ.
Decision accuracy
500 examples per task · development diagnosticsBanking77
- Reflex-1 95.0%
- OpenJev 90.6%
- Laya · native 44.4%
- Laya · expanded 55.4%
- Kev 82.4%
Emotion
- Reflex-1 92.2%
- OpenJev 53.0%
- Laya · native 57.6%
- Laya · expanded 57.6%
- Kev 54.4%
AG News
- Reflex-1 91.2%
- OpenJev 77.4%
- Laya · native 91.2%
- Laya · expanded 91.2%
- Kev 87.2%
SST-5
- Reflex-1 59.8%
- OpenJev 59.8%
- Laya · native 51.0%
- Laya · expanded 51.0%
- Kev 55.2%
BoolQ
- Reflex-1 86.0%
- OpenJev 86.8%
- Laya · native 71.2%
- Laya · expanded 71.2%
- Kev 82.6%
The task families were included in training and these subsets were inspected during development. Training exposure differs between models. Laya’s native budget truncates Banking77 options.
Full measurement table and protocol
| Task | reflex-1 | Laya, native | Laya, expanded | OpenJev | Kev |
|---|---|---|---|---|---|
| Banking77 | 95.0 | 44.4 | 55.4 | 90.6 | 82.4 |
| Emotion | 92.2 | 57.6 | 57.6 | 53.0 | 54.4 |
| AG News | 91.2 | 91.2 | 91.2 | 77.4 | 87.2 |
| SST-5 | 59.8 | 51.0 | 51.0 | 59.8 | 55.2 |
| BoolQ | 86.0 | 71.2 | 71.2 | 86.8 | 82.6 |
Laya SDK 0.3.20 is shown with native and expanded token budgets. The native Banking77 budget truncates the option set; the expanded budget retains all 77 labels. These are development diagnostics, not independent transfer evaluations.
These were 500-example diagnostics per task. The task families were included in training, and we had inspected the subsets during development. Training budgets and prior exposure differ between models. Accuracy was measured separately from latency, on more examples.
The general checkpoint completed 0 of 80 trials with our custom browser controller. The result exposes a substantial gap between familiar-task classification and autonomous control. Transfer to unfamiliar questions and probability calibration remain open evaluation problems.
Training and storage
The serving model has 420,778,370 parameters after merging adapters. ModernBERT encodes the state and questions; a frozen MiniLM encoder represents the choices. The shared head combines their representations.
| Resource | Recorded configuration |
|---|---|
| Serving parameters | 420,778,370 |
| Parameters trained in adaptation | 9,036,290 |
| Adaptation hardware | 1 × NVIDIA H200 |
| Main run | 6,000 updates · 75 min 7 sec |
| Extracted FP32 model on disk | ≈ 1.69 GB |
| Measured CPU process memory | 1.50 GiB peak RSS · separate 240-request replay |
| CPU compute threads | PyTorch intra-op/inter-op 1/1 · BLAS limit 1 |
The main adaptation run updated 9,036,290 parameters over 6,000 steps, using 290,657 records from eight task families. It took 75 minutes 7 seconds on one H200, including development screens and checkpoint saves. Base-model pretraining, data preparation, projection fitting and earlier experiments are additional costs.
The extracted FP32 model takes about 1.69 GB on disk. The separate CPU resource replay measured 1.50 GiB peak process RSS, including framework allocations, tokenizers, loading, warmup and inference. GPU memory usage remains unmeasured.
Model access
The public Hugging Face release contains the 1.68 GB FP32 weights, both tokenizers, and a self-contained Transformers implementation under Apache 2.0. No account, API key, or GitHub access is required. Use Python 3.12:
python -m pip install "torch==2.8.0" "transformers==5.17.0" "safetensors==0.8.0"
import torch
from transformers import AutoModel
torch.set_num_threads(1)
model = AutoModel.from_pretrained("gai-labs/reflex-1", trust_remote_code=True)
decision = model.predict(
state="The customer was charged twice for one card payment.",
question="Choose the matching issue.",
options=["duplicate charge", "lost card", "unknown fee", "cash withdrawal"],
)[0]
print(decision.choice)
Load the model once and reuse it. The model card documents batch inference, multiple questions, offline loading, and release hashes. The development repository remains private; the public package includes the code needed for inference.
The original training used RACE and SciQ, whose source terms include noncommercial restrictions. The model card documents their use and pinned revisions. The provenance audit does not establish how those terms apply to trained weights.
Recordings and reproducibility
Use the immutable October 3 Hub revision to reproduce these recordings. The charts above retain the measured outcomes, including retention regressions. Each recording identifies its environment, input format and playback speed.
The current release has a separate evaluation linked from its model card. These recordings and studies retain their original results.
Also in research
Samsara-VLA
Two compact robot policies for language-conditioned manipulation. Watch them pick, place, open and complete multi-step tasks.
VIO-lens-2
Research in progress.
Correspondence
Write to us
Contact us about evaluations, model access or a research collaboration.