EN
Contact
Menu
Journal

ANNOUNCEMENT

Decision models

Introducing Reflex-1

A 421M-parameter model for tool routing, intent classification and action selection, with public weights and measured CPU performance.

Good AI Labs 5 min read

Reflex-1 · Original Chromium43.6 SEC SILENT VIDEO
Reflex-1 · Chromium Dino. Reflex-1 in the native Chromium Dino game, using structured visible engine geometry. Development snapshot from 3 October 2026. This complete independent-course recording ends in collision at 42.309 seconds (score 535); it does not complete the 60-second budget. The browser keeps running during inference. Normal-speed browser recording; game clock continues during inference.

One pass, many possible decisions

Reflex-1 is a 421M-parameter model for decisions that end in a choice. Give it a textual state, a question, and candidate answers; it scores them in a single forward pass. Candidates can change with every request.

The same interface can route a support request, select a tool, or choose an action. A 28-layer context encoder and a six-layer candidate encoder feed a shared scoring head. The application supplies the choices and executes the result.

Reflex-1 / Decision model

From a question to a choice.

  1. 01 / Supply

    Define the decision.

    Your application supplies the state, a question and the choices it can act on. Those choices can change with every request: support queues, available tools or permitted actions.

    State + question

    The customer was charged twice.
    Which issue matches?

    Duplicate chargeLost cardUnknown feeCash withdrawal
    Context + questionModernBERT
    Each choiceMiniLM
    Shared scoring headChoice probabilities
    ApplicationPermission → arguments → action
    Illustrative support request. The application defines the candidate set.
  2. 02 / Encode

    Represent the context and the candidates.

    An adapted ModernBERT encoder reads the state and question. A frozen MiniLM encoder represents each supplied choice. Several questions can share a packed state.

    State + question

    The customer was charged twice.
    Which issue matches?

    Duplicate chargeLost cardUnknown feeCash withdrawal
    Context + questionModernBERT
    Each choiceMiniLM
    Shared scoring headChoice probabilities
    ApplicationPermission → arguments → action
    Two encoders feed one shared scoring head. This is a schematic of the preview architecture.
  3. 03 / Score

    Score every supplied choice in one pass.

    The shared head combines context and candidate representations, then returns a probability distribution over the supplied choices. New decision types still need accuracy and calibration tests.

    State + question

    The customer was charged twice.
    Which issue matches?

    Duplicate chargeLost cardUnknown feeCash withdrawal
    Context + questionModernBERT
    Each choiceMiniLM
    Shared scoring headChoice probabilities
    ApplicationPermission → arguments → action
    One score per supplied candidate, normalized to a probability distribution.
  4. 04 / Act

    Let the application carry out the decision.

    The application decides which actions are permitted, constructs any tool arguments and executes the selected action. Reflex-1 supplies the scores; the surrounding system owns the control loop.

    State + question

    The customer was charged twice.
    Which issue matches?

    Duplicate chargeLost cardUnknown feeCash withdrawal
    Context + questionModernBERT
    Each choiceMiniLM
    Shared scoring headChoice probabilities
    ApplicationPermission → arguments → action
    The application boundary matters: scoring a tool does not execute it.

Architecture of the public development preview. The support request illustrates the input format.

Recorded examples

The Dino recording above and the examples below use the October 3 snapshot, which differs from the default public weights. Run conditions and outcomes appear in the captions.

Reflex-1 · Native ViZDoom36.5 SEC SILENT VIDEO
Reflex-1 · ViZDoom. Reflex-1 in native ViZDoom 1.3.1, Defend the Center at skill 5. Development snapshot from 3 October 2026; predeclared recording seed 510001. The complete episode ends in death after 34.514 of the 45 allotted simulation seconds, with 30 kills. The model receives visible engine labels and health/ammo, not pixels. Playback follows recorded simulation time and excludes inference delays. Playback follows simulation time, excluding inference delays; original terminal hold retained.
Reflex-1 · Synthetic WebGym30.1 SEC · 4× SILENT VIDEO
Restaurant reservation · task completed. October 3 development snapshot; 15 actions. Browser states captured in Chromium and replayed at 4× wall time, with the complete ending and terminal hold. Local synthetic task, with a shared Qwen3-1.7B text helper. Snapshot differs from default Hub weights.

View all recordings and comparisons.

CPU inference

The current checkpoint was measured on an Intel Xeon Platinum 8558 CPU in FP32, with one request at a time. Each thread setting covers 480 calls over 120 inputs in two fresh processes. Timings include tokenization and inference.

Current CPU measurements

Intel Xeon Platinum 8558 · FP32 · batch size 1

One compute thread

Latency · lower is fasterMedian / p95
  1. AG News 1209.7 / 1385.0 ms
  2. SST-5 973.6 / 1160.3 ms
  3. Emotion 956.9 / 1165.8 ms
  4. Banking77 1101.7 / 1276.9 ms
  5. BoolQ 1600.8 / 2710.6 ms

Four compute threads

Latency · lower is fasterMedian / p95
  1. AG News 389.5 / 480.4 ms
  2. SST-5 320.6 / 404.2 ms
  3. Emotion 318.8 / 391.0 ms
  4. Banking77 375.2 / 452.1 ms
  5. BoolQ 486.7 / 782.2 ms

480 calls per thread setting over 120 inputs in two fresh processes. Tokenization and inference included; loading and warmup excluded. Shared host, evaluation runtime, no cross-request caches.

Peak process memory was 2.17 GiB, including loading, warmup, and inference. The runs used one or four PyTorch intra-op threads, one inter-op thread, and matching BLAS limits. Two and eleven total OS threads were observed, respectively, including runtime helpers.

Current measurements and evaluation. These shared-host timings use the evaluation runtime. Loading, warmup and cross-request caches are excluded. The research page retains earlier M4 and H200 measurements with their checkpoint dates.

Try Reflex-1

Public weights, tokenizers, and inference code are available on Hugging Face. No account or API key is required. With the dependencies installed:

import torch
from transformers import AutoModel

torch.set_num_threads(1)
model = AutoModel.from_pretrained("gai-labs/reflex-1", trust_remote_code=True)

decision = model.predict(
    state="The customer was charged twice for one card payment.",
    question="Choose the matching issue.",
    options=["duplicate charge", "lost card", "unknown fee", "cash withdrawal"],
)[0]
print(decision.choice)

See the model card for installation, runtime limits, and licensing, including the training-source terms.

MODEL DETAILS

Reflex-1

A 421M-parameter model for tool routing, intent classification and action selection. Public weights; runs on CPU or GPU.