EN
Contact
Menu
Journal

Systems

Budgeting a local decision service

Our reflex-1 timings start with the model loaded. A deployed service also has startup, queues and tool calls to account for.

The 27.5 ms H200 median for reflex-1 includes tokenization, scoring, output normalization and loopback HTTP. It starts with the model loaded and one request in flight. A service with a queue, cold starts or downstream tool calls will have a different response time.

The M4 CPU result measures an in-process request on a separate workload: 345.2 ms median for Banking77. Its timing also excludes loading. The benchmark report gives both protocols.

From a choice to an action

reflex-1 returns probabilities over choices supplied by the application. The application maps a choice to a tool and checks permission to execute it. Log the model version and the available choices with the selected answer and resulting action. This lets you distinguish a routing error from a tool that failed after being correctly selected.

Inputs and logs need the same access controls as the rest of the application. Local inference alone does not determine where its tools or retrieval services send data.

What to measure in deployment

Measure startup time and resident memory on the intended hardware. Then measure request latency under expected concurrency, including queue time and downstream work. Record the model settings and input sizes with the result.

The current reflex-1 report covers resident inference and model storage. Peak memory and complete application latency require these additional measurements.