EN
Contact
Menu
Journal

Model engineering

Adapting 9M parameters inside a 421M model

reflex-1's main adaptation run took 75 minutes on one H200. Where the trainable parameters sit and what that timing includes.

reflex-1’s main adaptation run updated 9,036,290 parameters. After the adapters were merged, the serving model contained 420,778,370. The trainable count was about 2.15% of the serving count.

Most parameters came from the pretrained encoders. ModernBERT represents the state and question; frozen MiniLM represents candidate answers. A fitted projection connects their representations. Rank-16 adapters and a shared decision head were trained in the main run.

The frozen weights still participate in inference. They occupy memory and perform computation even though adaptation does not update them.

The measured run

We used 290,657 records from eight task families and ran 6,000 updates. The loop took 75 minutes 7 seconds on one NVIDIA H200, including development screens and checkpoint saves.

That timing excludes base-model pretraining, downloads, data preparation, projection fitting and earlier experiments. The run produced a checkpoint with 95.0% Banking77 and 92.2% Emotion accuracy in familiar-task diagnostics. The performance report covers all five tasks and the compared checkpoints.

After merging

The extracted FP32 model takes about 1.69 GB on disk. Serving still needs memory for activations and framework allocations. The adapter fraction therefore describes the adaptation job, not the fraction of the full model needed at runtime.