# Introducing Reflex-1 · Good AI Labs > A 421M-parameter model for tool routing, intent classification and action selection, with public weights and measured CPU performance. URL: https://www.goodailabs.com/blog/reflex-1-fast-decisions/ Date: 2026-10-05 ## One pass, many possible decisions Reflex-1 is a 421M-parameter model for decisions that end in a choice. Give it a textual state, a question, and candidate answers; it scores them in a single forward pass. Candidates can change with every request. The same interface can route a support request, select a tool, or choose an action. A 28-layer context encoder and a six-layer candidate encoder feed a shared scoring head. The application supplies the choices and executes the result. From a question to a choice. Architecture of the public development preview. The support request illustrates the input format. 1. Define the decision. Your application supplies the state, a question and the choices it can act on. Those choices can change with every request: support queues, available tools or permitted actions. Figure: Illustrative support request. The application defines the candidate set. 2. Represent the context and the candidates. An adapted ModernBERT encoder reads the state and question. A frozen MiniLM encoder represents each supplied choice. Several questions can share a packed state. Figure: Two encoders feed one shared scoring head. This is a schematic of the preview architecture. 3. Score every supplied choice in one pass. The shared head combines context and candidate representations, then returns a probability distribution over the supplied choices. New decision types still need accuracy and calibration tests. Figure: One score per supplied candidate, normalized to a probability distribution. 4. Let the application carry out the decision. The application decides which actions are permitted, constructs any tool arguments and executes the selected action. Reflex-1 supplies the scores; the surrounding system owns the control loop. Figure: The application boundary matters: scoring a tool does not execute it. ## Recorded examples The Dino recording above and the examples below use the [October 3 snapshot](https://huggingface.co/gai-labs/reflex-1/tree/ae50bf81cddcaaac2aa18021d862433ca69da197), which differs from the default public weights. Run conditions and outcomes appear in the captions. Reflex-1 · ViZDoom. Reflex-1 in native ViZDoom 1.3.1, Defend the Center at skill 5. Development snapshot from 3 October 2026; predeclared recording seed 510001. The complete episode ends in death after 34.514 of the 45 allotted simulation seconds, with 30 kills. The model receives visible engine labels and health/ammo, not pixels. Playback follows recorded simulation time and excludes inference delays. Playback follows simulation time, excluding inference delays; original terminal hold retained. Recording or figure: https://www.goodailabs.com/media/reflex-1/preview/vizdoom-510001-clean.mp4 Restaurant reservation · task completed. October 3 development snapshot; 15 actions. Browser states captured in Chromium and replayed at 4× wall time, with the complete ending and terminal hold. Local synthetic task, with a shared Qwen3-1.7B text helper. Snapshot differs from default Hub weights. Recording or figure: https://www.goodailabs.com/media/reflex-1/preview/webgym-restaurant-960000-clean.mp4 [View all recordings and comparisons](/research/reflex-1/#examples). ## CPU inference The current checkpoint was measured on an Intel Xeon Platinum 8558 CPU in FP32, with one request at a time. Each thread setting covers 480 calls over 120 inputs in two fresh processes. Timings include tokenization and inference. Current CPU measurements. Intel Xeon Platinum 8558, FP32, batch size 1. 1 compute threads. Median / p95: - AG News: 1209.7 / 1385.0 ms. - SST-5: 973.6 / 1160.3 ms. - Emotion: 956.9 / 1165.8 ms. - Banking77: 1101.7 / 1276.9 ms. - BoolQ: 1600.8 / 2710.6 ms. 4 compute threads. Median / p95: - AG News: 389.5 / 480.4 ms. - SST-5: 320.6 / 404.2 ms. - Emotion: 318.8 / 391.0 ms. - Banking77: 375.2 / 452.1 ms. - BoolQ: 486.7 / 782.2 ms. 480 calls per thread setting over 120 inputs in two fresh processes. Tokenization and inference included; loading and warmup excluded. Shared host, evaluation runtime, no cross-request caches. Peak process memory was **2.17 GiB**, including loading, warmup, and inference. The runs used one or four PyTorch intra-op threads, one inter-op thread, and matching BLAS limits. Two and eleven total OS threads were observed, respectively, including runtime helpers. [Current measurements and evaluation](https://huggingface.co/gai-labs/reflex-1/blob/84c2cd49d31f73544e1e17539092056e741e7f03/evaluation.json). These shared-host timings use the evaluation runtime. Loading, warmup and cross-request caches are excluded. The research page retains [earlier M4 and H200 measurements](/research/reflex-1/#speed-and-resources) with their checkpoint dates. ## Try Reflex-1 Public weights, tokenizers, and inference code are available on [Hugging Face](https://huggingface.co/gai-labs/reflex-1). No account or API key is required. With the dependencies installed: ```python import torch from transformers import AutoModel torch.set_num_threads(1) model = AutoModel.from_pretrained("gai-labs/reflex-1", trust_remote_code=True) decision = model.predict( state="The customer was charged twice for one card payment.", question="Choose the matching issue.", options=["duplicate charge", "lost card", "unknown fee", "cash withdrawal"], )[0] print(decision.choice) ``` See the [model card](https://huggingface.co/gai-labs/reflex-1) for installation, runtime limits, and licensing, including the training-source terms. Contact: research@goodailabs.com