Systems
Choosing hardware for a local model
reflex-1 runs on a laptop CPU as well as an H200. The request, latency target and runtime memory determine which machine fits the application.
reflex-1 took 345.2 ms median to select among 77 intents on one M4 CPU thread. A separate H200 test returned recorded agent requests in 27.5 ms median, including local HTTP. A developer can run the model on either; the application’s latency target determines which result is relevant.
The requests differ between these tests. To choose hardware for a service, measure its actual input lengths and option counts at the expected concurrency. The reflex-1 report records the conditions for our measurements.
Disk, memory and adaptation
The FP32 reflex-1 model occupies about 1.69 GB on disk and contains 421M serving parameters. Its main adaptation run trained 9.04M parameters. Storage, inference and adaptation therefore have different budgets. Activations and framework allocations add to runtime memory, and the number of simultaneous requests can change the peak.
Samsara illustrates the training cost. Its 217M-parameter policy was trained on eight H200s. Near the end of the run, rank zero recorded 23.52 GiB peak allocated memory. That training measurement includes costs such as gradients and optimizer state; it does not specify inference requirements. The Samsara report gives the setup.
The rest of the application
Retrieval, tool execution and logging also need hardware. Keeping inference local controls where model inputs are processed, but each connected service has its own data path and resource use.
Where inference runs
An external endpointApplication context is sent to a model running outside the site
Your own hardwareApplication context is processed by a model deployed at the site
Measure the full request from arrival to completed action. Model latency is one part of that interval; queues and tool calls may account for more.