Decision models
Introducing Reflex-1
A 421M-parameter model for tool routing, intent classification and action selection, with public weights and measured CPU performance.
One pass, many possible decisions
Reflex-1 is a 421M-parameter model for decisions that end in a choice. Give it a textual state, a question, and candidate answers; it scores them in a single forward pass. Candidates can change with every request.
The same interface can route a support request, select a tool, or choose an action. A 28-layer context encoder and a six-layer candidate encoder feed a shared scoring head. The application supplies the choices and executes the result.
Reflex-1 / Decision model
From a question to a choice.
01 / Supply
Define the decision.
Your application supplies the state, a question and the choices it can act on. Those choices can change with every request: support queues, available tools or permitted actions.
State + question
The customer was charged twice.
Which issue matches?Context + questionModernBERTEach choiceMiniLMShared scoring headChoice probabilitiesApplicationPermission → arguments → actionIllustrative support request. The application defines the candidate set. 02 / Encode
Represent the context and the candidates.
An adapted ModernBERT encoder reads the state and question. A frozen MiniLM encoder represents each supplied choice. Several questions can share a packed state.
State + question
The customer was charged twice.
Which issue matches?Context + questionModernBERTEach choiceMiniLMShared scoring headChoice probabilitiesApplicationPermission → arguments → actionTwo encoders feed one shared scoring head. This is a schematic of the preview architecture. 03 / Score
Score every supplied choice in one pass.
The shared head combines context and candidate representations, then returns a probability distribution over the supplied choices. New decision types still need accuracy and calibration tests.
State + question
The customer was charged twice.
Which issue matches?Context + questionModernBERTEach choiceMiniLMShared scoring headChoice probabilitiesApplicationPermission → arguments → actionOne score per supplied candidate, normalized to a probability distribution. 04 / Act
Let the application carry out the decision.
The application decides which actions are permitted, constructs any tool arguments and executes the selected action. Reflex-1 supplies the scores; the surrounding system owns the control loop.
State + question
The customer was charged twice.
Which issue matches?Context + questionModernBERTEach choiceMiniLMShared scoring headChoice probabilitiesApplicationPermission → arguments → actionThe application boundary matters: scoring a tool does not execute it.
Architecture of the public development preview. The support request illustrates the input format.
Recorded examples
The Dino recording above and the examples below use the October 3 snapshot, which differs from the default public weights. Run conditions and outcomes appear in the captions.
View all recordings and comparisons.
CPU inference
The current checkpoint was measured on an Intel Xeon Platinum 8558 CPU in FP32, with one request at a time. Each thread setting covers 480 calls over 120 inputs in two fresh processes. Timings include tokenization and inference.
Current CPU measurements
Intel Xeon Platinum 8558 · FP32 · batch size 1One compute thread
- AG News 1209.7 / 1385.0 ms
- SST-5 973.6 / 1160.3 ms
- Emotion 956.9 / 1165.8 ms
- Banking77 1101.7 / 1276.9 ms
- BoolQ 1600.8 / 2710.6 ms
Four compute threads
- AG News 389.5 / 480.4 ms
- SST-5 320.6 / 404.2 ms
- Emotion 318.8 / 391.0 ms
- Banking77 375.2 / 452.1 ms
- BoolQ 486.7 / 782.2 ms
480 calls per thread setting over 120 inputs in two fresh processes. Tokenization and inference included; loading and warmup excluded. Shared host, evaluation runtime, no cross-request caches.
Peak process memory was 2.17 GiB, including loading, warmup, and inference. The runs used one or four PyTorch intra-op threads, one inter-op thread, and matching BLAS limits. Two and eleven total OS threads were observed, respectively, including runtime helpers.
Current measurements and evaluation. These shared-host timings use the evaluation runtime. Loading, warmup and cross-request caches are excluded. The research page retains earlier M4 and H200 measurements with their checkpoint dates.
Try Reflex-1
Public weights, tokenizers, and inference code are available on Hugging Face. No account or API key is required. With the dependencies installed:
import torch
from transformers import AutoModel
torch.set_num_threads(1)
model = AutoModel.from_pretrained("gai-labs/reflex-1", trust_remote_code=True)
decision = model.predict(
state="The customer was charged twice for one card payment.",
question="Choose the matching issue.",
options=["duplicate charge", "lost card", "unknown fee", "cash withdrawal"],
)[0]
print(decision.choice)
See the model card for installation, runtime limits, and licensing, including the training-source terms.
MODEL DETAILS
Reflex-1
A 421M-parameter model for tool routing, intent classification and action selection. Public weights; runs on CPU or GPU.