# Reflex-1 · Good AI Labs > A 421M-parameter decision model for agents. One forward pass, on CPU or GPU. URL: https://www.goodailabs.com/research/reflex-1/ Date: 2026-10-05 Status: Fine-tuned weights available on Hugging Face. Model card: https://huggingface.co/gai-labs/reflex-1 Software release: 1.0.0 ## Overview Many agent decisions end with a choice: which queue receives a request, which tool fits a task, or which action comes next. Reflex-1 scores the choices your application supplies in a single forward pass. It returns a probability distribution that the application can use to make that decision. The research follows the decision-model direction explored by [Jev](https://docs.typesafe.ai/introduction), with a distinct architecture and supervised training procedure. A 28-layer context encoder and a six-layer candidate encoder feed a shared scoring head. The current checkpoint updates all 420,778,370 parameters. The latest fine-tuned checkpoint is the default public download. Its model card contains current evaluation results and CPU measurements. The dated studies and recordings below retain their original checkpoint identities. Reflex-1 — Decision model Public release · runtime 1.0.0 Use: Choose among options supplied by an application: labels, routes or possible actions. Input → output: State, question and 1–255 choices → a choice distribution in one forward pass. Model: 421M parameters. A 28-layer context encoder, a 6-layer candidate encoder and a shared scoring head. Run locally: CPU or GPU · FP32. The current CPU replay measured 2.17 GiB peak process memory with one or four compute threads. Evaluation: 92.93% on a 495-case Banking77 reserved subset; 21/30 native MiniWoB++ episodes. Task families are represented in training. Access: Public weights, tokenizers and inference code under Apache 2.0. English text; runtime 1.0.0. The latest fine-tuned checkpoint is the default download. The dated studies and recordings below describe earlier checkpoints. The application supplies the choices and controls execution. Weights & inference: https://huggingface.co/gai-labs/reflex-1 Release results: https://www.goodailabs.com/research/reflex-1/#current-results ## Current results The October 5 release completed **21/30 native MiniWoB++ browser episodes** with a fixed controller. It scored **92.93% on 495 Banking77 examples**. These are reserved subsets from task families represented in training, not full benchmark scores. October 5 checkpoint — reserved classification subsets. - AG News: 465/500 correct (93.00%). - SST-5: 299/500 correct (59.80%). - Emotion: 458/500 correct (91.60%). - Banking77: 460/495 correct (92.93%). - BoolQ: 171/202 correct (84.65%). - MNLI: 436/500 correct (87.20%). - RACE: 287/500 correct (57.40%). - SciQ: 123/128 correct (96.09%). On an Intel Xeon Platinum 8558 CPU, Banking77 inference took **375.2 ms median** and **452.1 ms p95** with four compute threads. Peak process memory was **2.17 GiB**, including loading, warmup and inference. Current CPU measurements. Intel Xeon Platinum 8558, FP32, batch size 1. 1 compute threads. Median / p95: - AG News: 1209.7 / 1385.0 ms. - SST-5: 973.6 / 1160.3 ms. - Emotion: 956.9 / 1165.8 ms. - Banking77: 1101.7 / 1276.9 ms. - BoolQ: 1600.8 / 2710.6 ms. 4 compute threads. Median / p95: - AG News: 389.5 / 480.4 ms. - SST-5: 320.6 / 404.2 ms. - Emotion: 318.8 / 391.0 ms. - Banking77: 375.2 / 452.1 ms. - BoolQ: 486.7 / 782.2 ms. 480 calls per thread setting over 120 inputs in two fresh processes. Tokenization and inference included; loading and warmup excluded. Shared host, evaluation runtime, no cross-request caches. The full evaluation is still incomplete. BoolQ, RACE and the public JevBench diagnostic regressed from the previous checkpoint; native Dino reached the 60-second target in **0/16 attempts**. Current GPU latency has not been measured. The [published evaluation](https://huggingface.co/gai-labs/reflex-1/blob/84c2cd49d31f73544e1e17539092056e741e7f03/evaluation.json) records the checkpoint, conditions and remaining gaps. ## How it works Candidate answers are supplied with each request. The model supports 1–255 choices per question, within its token limits, and several questions can share a packed state. A support application can supply its own queues; an agent can supply descriptions of the tools it has available. New decision types need their own accuracy tests. The model scores the supplied options; the application constructs tool arguments and decides which actions are permitted. State text is capped at 512 context-tokenizer tokens and can be shortened further by the request budget. Each choice accepts up to 512 option-tokenizer tokens. Inspect `state_truncated` when context may have been shortened. From a question to a choice. Architecture of the public development preview. The support request illustrates the input format. 1. Define the decision. Your application supplies the state, a question and the choices it can act on. Those choices can change with every request: support queues, available tools or permitted actions. Figure: Illustrative support request. The application defines the candidate set. 2. Represent the context and the candidates. An adapted ModernBERT encoder reads the state and question. A frozen MiniLM encoder represents each supplied choice. Several questions can share a packed state. Figure: Two encoders feed one shared scoring head. This is a schematic of the preview architecture. 3. Score every supplied choice in one pass. The shared head combines context and candidate representations, then returns a probability distribution over the supplied choices. New decision types still need accuracy and calibration tests. Figure: One score per supplied candidate, normalized to a probability distribution. 4. Let the application carry out the decision. The application decides which actions are permitted, constructs any tool arguments and executes the selected action. Reflex-1 supplies the scores; the surrounding system owns the control loop. Figure: The application boundary matters: scoring a tool does not execute it. ## Examples The examples and decision-quality comparisons below use the **October 3 snapshot**. Speed and resource measurements are labelled separately. Each recording keeps its observed ending, including failure. ### Chromium Dino This recording uses Chromium's original `chrome://dino` game. Reflex-1 receives the currently visible obstacle geometry and chooses run, jump or duck. Native keyboard events execute those choices while the game clock continues during inference. The video shows actual browser frames at normal speed. Reflex-1 · Chromium Dino. Reflex-1 in the native Chromium Dino game, using structured visible engine geometry. Development snapshot from 3 October 2026. This complete independent-course recording ends in collision at 42.309 seconds (score 535); it does not complete the 60-second budget. The browser keeps running during inference. Normal-speed browser recording; game clock continues during inference. Recording or figure: https://www.goodailabs.com/media/reflex-1/preview/chromium-dino-002-clean.mp4 The two attempts with this snapshot ended in collisions after **42.31 seconds** and **30.41 seconds**; neither reached the 60-second limit. The longer attempt is shown first. These are independent random courses, so they do not support a paired model comparison. The complete second attempt remains available below. Reflex-1 · Chromium Dino. Reflex-1 in the native Chromium Dino game, using structured visible engine geometry. Development snapshot from 3 October 2026. This complete independent-course recording ends in collision at 30.410 seconds (score 351); it does not complete the 60-second budget. The browser keeps running during inference. Normal-speed browser recording; game clock continues during inference. Recording or figure: https://www.goodailabs.com/media/reflex-1/preview/chromium-dino-003-clean.mp4 ### ViZDoom The game is native ViZDoom's **Defend the Center** scenario at skill 5. Reflex-1 chooses actions from visible object bounds, health and ammunition. This is a structured-state controller; it does not infer actions from pixels. The recording below is the originally predeclared example, shown as a single Reflex-1 view. Reflex-1 · ViZDoom. Reflex-1 in native ViZDoom 1.3.1, Defend the Center at skill 5. Development snapshot from 3 October 2026; predeclared recording seed 510001. The complete episode ends in death after 34.514 of the 45 allotted simulation seconds, with 30 kills. The model receives visible engine labels and health/ammo, not pixels. Playback follows recorded simulation time and excludes inference delays. Playback follows simulation time, excluding inference delays; original terminal hold retained. Recording or figure: https://www.goodailabs.com/media/reflex-1/preview/vizdoom-510001-clean.mp4 Across all 32 reserved episodes, this snapshot averaged **18.88 kills** and **22.47 seconds** of survival. Every episode ended before the 45-second limit. The chart retains every evaluated external model and the full episode counts. ViZDoom · kills and survival limits. 3 October 2026 · development snapshot 70a843878214 Reflex-1 denotes the October 3 development weights shown here, not the default download or a final release. Original ViZDoom engine, structured visible-state inputs and raw argmax actions; 45-second episode limit. Every declared outcome is retained. Mean intervals resample paired seeds 10,000 times. Reserved episode seeds do not establish unseen structured states. This named adapter/environment comparison is not an official benchmark score or a pixel-input policy evaluation. Reflex-1 averages 18.88 kills and 22.47 seconds of survival, with 0/32 full-duration completions. Laya, OpenJev and Kev each average 1.375 kills and also complete 0/32. The reference models were used as shipped; training histories differ. Defend the Center · skill 5 — Kills · higher is better - Reflex-1: 18.9; 95% CI 16.6–21.2; 0/32 reached 45 s · 32 deaths · 0 errors · 22.47 s mean survival - Laya: 1.4; 95% CI 1.2–1.6; 0/32 reached 45 s · 32 deaths · 0 errors · 4.90 s mean survival - OpenJev · DeBERTa: 1.4; 95% CI 1.2–1.6; 0/32 reached 45 s · 32 deaths · 0 errors · 4.90 s mean survival - Kev-0.8B: 1.4; 95% CI 1.2–1.6; 0/32 reached 45 s · 32 deaths · 0 errors · 4.90 s mean survival Protocol and metrics: https://www.goodailabs.com/research/reflex-1/preview/figures/doom.json Chart data: https://www.goodailabs.com/research/reflex-1/preview/figures/doom.csv ### Browser workflows These are locally hosted synthetic WebGym sites distributed with Laya Browser. The browser states were captured in real Chromium, and the original task checker decides success. Reflex-1 selects actions and targets; a shared Qwen3-1.7B helper supplies typed text and dropdown values. The outcome measures that complete system. Restaurant reservation · task completed. October 3 development snapshot; 15 actions. Browser states captured in Chromium and replayed at 4× wall time, with the complete ending and terminal hold. Local synthetic task, with a shared Qwen3-1.7B text helper. Snapshot differs from default Hub weights. Recording or figure: https://www.goodailabs.com/media/reflex-1/preview/webgym-restaurant-960000-clean.mp4 Shopping task · goal not completed. October 3 development snapshot; 4 actions. Browser states captured in Chromium and replayed at 4× wall time, with the complete ending and terminal hold. Local synthetic task, with a shared Qwen3-1.7B text helper. Snapshot differs from default Hub weights. Recording or figure: https://www.goodailabs.com/media/reflex-1/preview/webgym-shop-960000-clean.mp4 The restaurant task succeeds; the shopping task fails. Both use the same predeclared task seed and retain the recorded ending. Playback is **4×** the original wall time, replaying captured browser states with the terminal hold preserved. These examples are synthetic-site development results, not MiniWoB++, WebArena or live-site scores. Browser completion · synthetic WebGym. 3 October 2026 · development snapshot 70a843878214 Reflex-1 denotes the October 3 development weights shown here, not the default download or a final release. Real Chromium capture of local synthetic WebGym sites. Every arm uses the same executor, Qwen3-1.7B text helper and 40-step limit. Reflex-1 uses BF16 autocast and 1,024 state / 16,384 request tokens; other serving conditions are preserved in the download. All 35 attempts per arm, request failures and helper errors remain included. This panel is separate from MiniWoB++, WebArena and live-site evaluation. The source site's final-state check records 13/35 completions for Reflex-1 and 17/35 for Laya Browser. The specialist remains ahead; the other three external arms are also shown. — Completed episodes (%) · higher is better - Reflex-1: 13/35; 485 decision calls (0 errors) · 102 helper calls (0 HTTP / 5 contract errors) - Laya Browser: 17/35; 505 decision calls (0 errors) · 100 helper calls (0 HTTP / 7 contract errors) - Laya: 0/35; 56 decision calls (0 errors) · 12 helper calls (0 HTTP / 0 contract errors) - OpenJev · DeBERTa: 0/35; 35 decision calls (24 errors) · 0 helper calls (0 HTTP / 0 contract errors) - Kev-0.8B: 1/35; 196 decision calls (0 errors) · 82 helper calls (0 HTTP / 0 contract errors) Protocol and metrics: https://www.goodailabs.com/research/reflex-1/preview/figures/browser.json Chart data: https://www.goodailabs.com/research/reflex-1/preview/figures/browser.csv Reflex-1 completed **13/35** tasks in this study; Laya Browser completed **17/35**. All attempts, failures and model-request errors remain in the denominators. A later pilot through the official MiniWoB++ environment completed **0/30** tasks with this snapshot, so synthetic-site results should not be read as success on that benchmark. ## Decision quality Typed Decisions covers customer support, invoice processing, security alerts and agent traces. The evaluation preserves original state groups and asks each model the same five judgments through a common choice interface. Reflex-1 answered **1,508/2,000** judgments correctly (**75.4%**). Structured decisions. 3 October 2026 · development snapshot 70a843878214 Reflex-1 denotes the October 3 development weights shown here, not the default download or a final release. 400 reserved source states, expanded to 2,000 judgments. Each workflow contains 500 judgments. Intervals resample source states, preserving the five related heads. Token limits: Reflex-1: state 1,024, request 4,096; Laya · native: request 512, instruction/choices 192; OpenJev · DeBERTa: state 256, request 512; Kev-0.8B: state 512, branch including state 2,048. Whole-cohort instruction/choice shortening (not additive across task charts): Laya · native 0/2,000. These counts are separate from the state cuts above. Agreement with the modal teacher choice through a common choice interface. These are adapted dataset results, not an official multihead leaderboard score. Customer support — Accuracy (%) · higher is better - Reflex-1: 76.6%; 95% CI 72.0–80.8%; 383/500 correct · 0 errors · 0 state cuts - Laya · native: 42.0%; 95% CI 38.2–45.8%; 210/500 correct · 0 errors · 39 state cuts - OpenJev · DeBERTa: 41.2%; 95% CI 37.8–44.0%; 206/500 correct · 0 errors · 350 state cuts - Kev-0.8B: 62.8%; 95% CI 58.2–67.4%; 314/500 correct · 0 errors · 10 state cuts Invoice processing — Accuracy (%) · higher is better - Reflex-1: 78.8%; 95% CI 75.0–82.4%; 394/500 correct · 0 errors · 0 state cuts - Laya · native: 38.4%; 95% CI 32.6–44.6%; 192/500 correct · 0 errors · 0 state cuts - OpenJev · DeBERTa: 40.6%; 95% CI 38.2–43.0%; 203/500 correct · 0 errors · 500 state cuts - Kev-0.8B: 50.4%; 95% CI 45.2–55.8%; 252/500 correct · 0 errors · 0 state cuts Security alerts — Accuracy (%) · higher is better - Reflex-1: 72.4%; 95% CI 68.4–76.2%; 362/500 correct · 0 errors · 0 state cuts - Laya · native: 33.4%; 95% CI 30.0–37.0%; 167/500 correct · 0 errors · 0 state cuts - OpenJev · DeBERTa: 42.0%; 95% CI 39.2–44.6%; 210/500 correct · 0 errors · 275 state cuts - Kev-0.8B: 47.2%; 95% CI 42.6–51.8%; 236/500 correct · 0 errors · 0 state cuts Agent traces — Accuracy (%) · higher is better - Reflex-1: 73.8%; 95% CI 69.0–78.4%; 369/500 correct · 0 errors · 0 state cuts - Laya · native: 38.6%; 95% CI 33.4–44.0%; 193/500 correct · 0 errors · 0 state cuts - OpenJev · DeBERTa: 38.4%; 95% CI 34.6–42.2%; 192/500 correct · 0 errors · 0 state cuts - Kev-0.8B: 35.8%; 95% CI 31.6–40.4%; 179/500 correct · 0 errors · 0 state cuts Protocol and metrics: https://www.goodailabs.com/research/reflex-1/preview/figures/typed.json Chart data: https://www.goodailabs.com/research/reflex-1/preview/figures/typed.csv On the reserved CLINC cohort, Reflex-1 answered **3,680/4,490** intent and out-of-scope requests correctly (**82.0%**). For unfamiliar-tool selection, it answered **72/87**, below the strongest external result of **82/87**. Both strengths and remaining gaps appear below. Intent routing. 3 October 2026 · development snapshot 70a843878214 Reflex-1 denotes the October 3 development weights shown here, not the default download or a final release. 4,490 reserved requests: 3,548 known intents and 942 out-of-scope requests. Native Laya and its separately declared expanded-context control are both shown. Token limits: Reflex-1: state 1,024, request 4,096; Laya · native: request 512, instruction/choices 192; Laya · expanded: request 2,048, instruction/choices 1,536; OpenJev · DeBERTa: state 256, request 512; Kev-0.8B: state 512, branch including state 2,048. Whole-cohort instruction/choice shortening (not additive across task charts): Laya · native 4,490/4,490; Laya · expanded 0/4,490. These counts are separate from the state cuts above. Reflex-1 correctly routes 3,108/3,548 known-intent requests and rejects 572/942 out-of-scope requests. It falsely rejects 64 known requests and accepts 370 out-of-scope requests. All external conditions are retained. Known and unfamiliar requests — Accuracy (%) · higher is better - Reflex-1: 82.0%; 95% CI 80.8–83.1%; 3,680/4,490 correct · 0 errors · 0 state cuts · 64 false rejections · 370 false acceptances - Laya · native: 31.7%; 95% CI 30.3–33.2%; 1,425/4,490 correct · 0 errors · 0 state cuts · 0 false rejections · 942 false acceptances - Laya · expanded: 22.7%; 95% CI 21.4–23.9%; 1,017/4,490 correct · 0 errors · 0 state cuts · 0 false rejections · 942 false acceptances - OpenJev · DeBERTa: 55.6%; 95% CI 54.1–57.0%; 2,496/4,490 correct · 0 errors · 0 state cuts · 5 false rejections · 939 false acceptances - Kev-0.8B: 51.2%; 95% CI 49.6–52.6%; 2,298/4,490 correct · 0 errors · 0 state cuts · 1 false rejections · 937 false acceptances Protocol and metrics: https://www.goodailabs.com/research/reflex-1/preview/figures/routing-clinc.json Chart data: https://www.goodailabs.com/research/reflex-1/preview/figures/routing-clinc.csv Selection among unfamiliar tools. 3 October 2026 · development snapshot 70a843878214 Reflex-1 denotes the October 3 development weights shown here, not the default download or a final release. 87 reserved requests with tool names and normalized states separated before training. This is a tool-choice adaptation; argument generation and execution are outside its scope. Token limits: Reflex-1: state 1,024, request 4,096; Laya · native: request 512, instruction/choices 192; Laya · expanded: request 2,048, instruction/choices 1,536; OpenJev · DeBERTa: state 256, request 512; Kev-0.8B: state 512, branch including state 2,048. Whole-cohort instruction/choice shortening (not additive across task charts): Laya · native 87/87; Laya · expanded 65/87. These counts are separate from the state cuts above. Reflex-1 selects 72/87 tools correctly. Laya reaches 80/87 in each declared condition, OpenJev 82/87 and Kev 81/87. This evaluates supplied tool choice, not argument generation or execution. Tool selection — Accuracy (%) · higher is better - Reflex-1: 82.8%; 95% CI 74.7–90.8%; 72/87 correct · 0 errors · 0 state cuts - Laya · native: 92.0%; 95% CI 86.2–97.7%; 80/87 correct · 0 errors · 14 state cuts - Laya · expanded: 92.0%; 95% CI 86.2–97.7%; 80/87 correct · 0 errors · 0 state cuts - OpenJev · DeBERTa: 94.3%; 95% CI 88.5–98.9%; 82/87 correct · 0 errors · 30 state cuts - Kev-0.8B: 93.1%; 95% CI 87.4–97.7%; 81/87 correct · 0 errors · 1 state cuts Protocol and metrics: https://www.goodailabs.com/research/reflex-1/preview/figures/routing-hermes.json Chart data: https://www.goodailabs.com/research/reflex-1/preview/figures/routing-hermes.csv These are common-choice adaptations with declared cohorts, not official leaderboard submissions. External models were evaluated as shipped; they did not receive Reflex-1's task-specific training. The figures retain the original uncertainty intervals, token budgets and truncation counts. ## Speed and resources The measurements in this section describe the October 2 checkpoint. Current checkpoint measurements are available in the model card linked above. ### H200 request latency We replayed 120 distinct requests twice per model on a shared NVIDIA H200. They came from five game and browser environments. Every model received the same inputs, with one request in flight at a time. H200 request latency. H200 median and p95 request latency in milliseconds: Reflex-1 FP32, 27.5 and 30.5; Laya native BF16, 27.1 and 142.0; Laya FP32 replay, 40.2 and 44.4; public OpenJev FP32, 32.0 and 33.8. 120 identical inputs, each replayed twice, with one request in flight. Timing includes formatting, tokenization, inference, and loopback HTTP. Spans show median to p95. Precision differs as labeled; the FP32 Laya replay changed two decisions. Chart: https://www.goodailabs.com/img/reflex-1/latency-h200.64f08ce1108e707fe13c7cd6b12e4bc7095c26d45bc3f151003422bb7dbd2337.svg Data: https://www.goodailabs.com/img/reflex-1/latency-h200.05d322dd52c147f2153d8df6bc3858a028ce186fb5251203bbf709c805c8766a.csv H200 · identical requests · milliseconds · lower is better | Model | Compute | Median | p95 | | --- | --- | ---: | ---: | | reflex-1 | FP32 | 27.5 | 30.5 | | OpenJev | FP32 | 32.0 | 33.8 | | Laya (native) | BF16 | 27.1 | 142.0 | | Laya (FP32 replay) | FP32 | 40.2 | 44.4 | Identical recorded requests, two repeats each, from five agent environments. Tokenization, model scoring, output normalization and loopback HTTP included. Three warmups and loading excluded. No answer or option cache. Four PyTorch CPU threads per server. Laya native uses BF16 autocast; a separate FP32 replay is also included. OpenJev is the public DeBERTa checkpoint, not hosted Jev. These request timings do not measure task success or concurrent throughput. reflex-1's median was 27.5 ms and its p95 was 30.5 ms. It was faster than OpenJev in FP32 and Laya in the separate FP32 replay. Native Laya used BF16 autocast and had a slightly lower median. The shared machine and sample size limit what we can infer from close results. The timer includes tokenization, model scoring, output normalization and loopback HTTP. It excludes loading and three warmup requests. Answer and option caches were disabled. We did not measure concurrent serving throughput in this run. OpenJev here is the public DeBERTa checkpoint; it is not the hosted Jev service. ### M4 CPU latency The CPU test used a MacBook Air M4, FP32, and batch size one. Each model used one PyTorch intra-op thread, one inter-op thread, and a BLAS thread limit of one. Each task had 24 distinct cases, each repeated twice. Model order was balanced across eight execution blocks. M4 CPU request latency. Request latency by task for Reflex-1, Laya, public OpenJev, and Kev-0.8B on an M4 CPU. Dots mark medians and lines end at p95 on a logarithmic time axis. FP32, batch size one; PyTorch intra-op/inter-op threads 1/1 and BLAS thread limit 1. Original latency study: 24 inputs per task, each repeated twice. A separate 240-request Reflex-1 replay measured 1.50 GiB peak process RSS, including loading, warmup and inference. Per-model memory was not recorded in the original comparison. Spans show median to p95. Laya truncates Banking77 options; Kev uses unmerged reference CPU kernels. Chart: https://www.goodailabs.com/img/reflex-1/latency-m4.d1c3121935e3d021edb1eff967d1cc6dbe955fe274927777dbb76804ff141da2.svg Data: https://www.goodailabs.com/img/reflex-1/latency-m4.9c19403e21fcb9a3124b0f0d747b82cd61f840bbf96fd73e44ab8a52dbd99554.csv M4 CPU · FP32 · PyTorch intra-op/inter-op 1/1 · BLAS thread limit 1 · median / p95 in milliseconds | Task | reflex-1 | OpenJev | Laya | Kev | | --- | ---: | ---: | ---: | ---: | | Banking77 | 345.2 / 692.0 | 1332.6 / 2857.1 | 912.4 / 2649.9 | 4283.4 / 11073.7 | | Emotion | 256.2 / 610.4 | 381.8 / 1029.5 | 240.1 / 427.7 | 1121.9 / 4422.3 | | AG News | 321.7 / 578.7 | 424.5 / 934.0 | 324.1 / 760.5 | 1356.5 / 4997.2 | | SST-5 | 299.1 / 631.0 | 360.5 / 894.5 | 250.7 / 548.5 | 902.2 / 4480.5 | | BoolQ | 463.4 / 1329.1 | 451.9 / 1398.9 | 440.4 / 1587.9 | 1775.3 / 6548.0 | Complete request: construction, tokenization, model inference, output normalization. One resident model at a time; no cross-request caches. Loading, warmup and cache clearing excluded. Eight balanced execution blocks; background apps active. Native runtime budgets. Laya truncates Banking77 options at its native budget; expanded-budget results are reported separately. Kev uses its CPU reference kernels and unmerged adapters. Reflex-1 measured memory: 1.50 GiB peak process RSS. Separate Reflex-1 resource replay of the original 120 inputs, each scored twice. Whole-process peak RSS includes imports, model loading, warmup and inference. The original four-model latency comparison did not record per-model memory. Recorded 2026-10-02 UTC. Resource measurements and protocol: https://huggingface.co/Goodailabs/reflex-1/resolve/main/assets/cpu-resources.json Banking77 asks the model to choose among 77 intents. reflex-1 took 345.2 ms median; OpenJev took 1,332.6 ms in the same run. Across the five tasks, reflex-1's medians ranged from 256.2 to 463.4 ms. Laya was faster on several smaller-choice tasks. Its native Banking77 budget truncated the options, so its timing covers a reduced candidate set. Kev used CPU reference kernels with unmerged adapters. Background laptop activity contributed to the wider tail latencies. This test includes request construction through normalized output, with the model already loaded. The CPU and GPU tests use different requests; their timings do not give a direct CPU-to-GPU speedup. ### Familiar-task accuracy reflex-1 scored 95.0% on Banking77 and 92.2% on Emotion. It led the compared checkpoints on both, tied the best score on AG News and SST-5, and scored slightly below OpenJev on BoolQ. Five-task classification accuracy. Accuracy on five 500-example diagnostic subsets. Reflex-1 scores 91.2% on AG News, 59.8% on SST-5, 92.2% on Emotion, 95.0% on Banking77, and 86.0% on BoolQ. The figure also shows native and expanded-budget Laya, public OpenJev, and Kev-0.8B. Each model received the same 500 examples per task. These task families were included in Reflex training, and the subsets were inspected during development. Training exposure differs across models. Expanded-budget Laya retains all 77 Banking77 labels. Chart: https://www.goodailabs.com/img/reflex-1/accuracy.3128aac64fb941e8d8f3db6cafac9e491ce2f2b7164f20d1add3f928fa797b8f.svg Data: https://www.goodailabs.com/img/reflex-1/accuracy.6f5d568dfe4adc41ad1ab8e26aa63364f71bac0250c9fff82d25dc12382b4704.csv Accuracy (%) · 500 examples per task · development diagnostics | Task | reflex-1 | Laya, native | Laya, expanded | OpenJev | Kev | | --- | ---: | ---: | ---: | ---: | ---: | | Banking77 | 95 | 44.4 | 55.4 | 90.6 | 82.4 | | Emotion | 92.2 | 57.6 | 57.6 | 53 | 54.4 | | AG News | 91.2 | 91.2 | 91.2 | 77.4 | 87.2 | | SST-5 | 59.8 | 51 | 51 | 59.8 | 55.2 | | BoolQ | 86 | 71.2 | 71.2 | 86.8 | 82.6 | Previously inspected diagnostics on task families included in training. Training budgets and prior exposure differ between models. reflex-1 accuracy was measured in GPU FP32 with TF32 off; baselines used CPU FP32. CPU output parity was checked on timing cases. Laya SDK 0.3.20 is shown with native and expanded token budgets. The native Banking77 budget truncates the option set; the expanded budget retains all 77 labels. These are development diagnostics, not independent transfer evaluations. These were 500-example diagnostics per task. The task families were included in training, and we had inspected the subsets during development. Training budgets and prior exposure differ between models. Accuracy was measured separately from latency, on more examples. The general checkpoint completed 0 of 80 trials with our custom browser controller. The result exposes a substantial gap between familiar-task classification and autonomous control. Transfer to unfamiliar questions and probability calibration remain open evaluation problems. ### Training and storage The serving model has 420,778,370 parameters after merging adapters. ModernBERT encodes the state and questions; a frozen MiniLM encoder represents the choices. The shared head combines their representations. - Serving parameters: 420778370. - Parameters trained in adaptation: 9036290. - Main adaptation run: 6,000 updates, 75 min 7 sec on 1 × NVIDIA H200. - Extracted FP32 model on disk: approximately 1.69 GB. - Measured CPU memory: 1.50 GiB peak process RSS in a separate 240-request replay. - CPU compute threads: PyTorch intra-op/inter-op 1/1; BLAS limit 1. The main adaptation run updated 9,036,290 parameters over 6,000 steps, using 290,657 records from eight task families. It took 75 minutes 7 seconds on one H200, including development screens and checkpoint saves. Base-model pretraining, data preparation, projection fitting and earlier experiments are additional costs. The extracted FP32 model takes about 1.69 GB on disk. The separate CPU resource replay measured 1.50 GiB peak process RSS, including framework allocations, tokenizers, loading, warmup and inference. GPU memory usage remains unmeasured. ## Model access The [public Hugging Face release](https://huggingface.co/gai-labs/reflex-1) contains the 1.68 GB FP32 weights, both tokenizers, and a self-contained Transformers implementation under Apache 2.0. No account, API key, or GitHub access is required. Use Python 3.12: ```bash python -m pip install "torch==2.8.0" "transformers==5.17.0" "safetensors==0.8.0" ``` ```python import torch from transformers import AutoModel torch.set_num_threads(1) model = AutoModel.from_pretrained("gai-labs/reflex-1", trust_remote_code=True) decision = model.predict( state="The customer was charged twice for one card payment.", question="Choose the matching issue.", options=["duplicate charge", "lost card", "unknown fee", "cash withdrawal"], )[0] print(decision.choice) ``` Load the model once and reuse it. The model card documents batch inference, multiple questions, offline loading, and release hashes. The development repository remains private; the public package includes the code needed for inference. The original training used RACE and SciQ, whose source terms include noncommercial restrictions. The [model card](https://huggingface.co/gai-labs/reflex-1) documents their use and pinned revisions. The provenance audit does not establish how those terms apply to trained weights. ## Recordings and reproducibility Use the [immutable October 3 Hub revision](https://huggingface.co/gai-labs/reflex-1/tree/ae50bf81cddcaaac2aa18021d862433ca69da197) to reproduce these recordings. The charts above retain the measured outcomes, including retention regressions. Each recording identifies its environment, input format and playback speed. The current release has a separate evaluation linked from its model card. These recordings and studies retain their original results. Measurement summary: https://www.goodailabs.com/research/reflex-1/evaluation.json Contact: research@goodailabs.com