EN
Contact
Menu
Journal

Systems

Choosing hardware for a local model

reflex-1 runs on a laptop CPU as well as an H200. The request, latency target and runtime memory determine which machine fits the application.

reflex-1 took 345.2 ms median to select among 77 intents on one M4 CPU thread. A separate H200 test returned recorded agent requests in 27.5 ms median, including local HTTP. A developer can run the model on either; the application’s latency target determines which result is relevant.

The requests differ between these tests. To choose hardware for a service, measure its actual input lengths and option counts at the expected concurrency. The reflex-1 report records the conditions for our measurements.

Disk, memory and adaptation

The FP32 reflex-1 model occupies about 1.69 GB on disk and contains 421M serving parameters. Its main adaptation run trained 9.04M parameters. Storage, inference and adaptation therefore have different budgets. Activations and framework allocations add to runtime memory, and the number of simultaneous requests can change the peak.

Samsara illustrates the training cost. Its 217M-parameter policy was trained on eight H200s. Near the end of the run, rank zero recorded 23.52 GiB peak allocated memory. That training measurement includes costs such as gradients and optimizer state; it does not specify inference requirements. The Samsara report gives the setup.

The rest of the application

Retrieval, tool execution and logging also need hardware. Keeping inference local controls where model inputs are processed, but each connected service has its own data path and resource use.

Where inference runs

01 / Remote

An external endpointApplication context is sent to a model running outside the site

Request → network → modelThe provider operates the inference hardware
02 / Local

Your own hardwareApplication context is processed by a model deployed at the site

Request → local modelThe deployment supplies its compute and capacity
Remote inference sends inputs to an external endpoint. Local inference runs the model on the application's own hardware.

Measure the full request from arrival to completed action. Model latency is one part of that interval; queues and tool calls may account for more.