The 125B badge is real. It is not a memory map.
Qwen lists Qwen3.8-Flash-Next as a 125-billion-parameter language model with 6 billion parameters activated for each token. The model card also names 51 billion n-gram embedding parameters, a 4-billion-parameter multi-token prediction layer, 512 experts, and 11 active experts per mixture layer.[2]
Those numbers answer different questions. Total parameters describe stored model capacity. Activated parameters describe one compute path. Neither number says how many bytes sit in graphics memory, system memory, or storage after quantization.
Active parameters explain work per token. They do not erase the weights that still need a home.
Strata turns placement into the product.
Strata's public design keeps frequently used experts on the graphics card. It keeps the full expert set in system memory and puts a large lookup table on storage. The project says its normal route needs at least 12 GB of graphics memory, 32 GB of system memory, and about 80 GB of free disk. It says setup downloads about 70 GB and loads 35 to 55 GB into system memory.[1]
The project publishes measured prompt-processing and generation rates for named quantizations on an RTX 5070 and an RX 9070 XT. It also records the engine versions, a 32K-token prompt, and a 4K-token answer. Those details make the table useful. The headline alone does not preserve them.[1]
Prompt speed and answer speed need separate dials.
llama.cpp's benchmark tool defines prompt processing, text generation, and a combined test as separate jobs. It repeats tests, reports an average and standard deviation, accepts a context depth, and records placement controls such as GPU layers and CPU mixture-of-experts layers. Its documentation also says benchmark timing excludes tokenization and sampling.[3]
A local coding run needs one more timing boundary. Record model load, first response, prompt processing, text generation, and wall-clock task completion. A fast decode cannot repay a long model load if the process restarts for every task. A fast prompt pass cannot prove that edits, tool calls, or tests finish sooner.
Quantization is part of the model identity.
Strata offers several compressed variants. Its README says smaller variants run faster while larger variants keep more quality. The Coder option removes half of the experts to fit 32 GB of system memory. The project also warns that a larger 4-bit option can read weights from storage when system memory is tight, which cuts generation speed.[1]
Write the exact model family, quantization, engine revision, draft model, context size, cache format, and placement beside every result. Without that row, two runs with the same 125B badge may test different artifacts and different routes through the machine.
The artifact exists. This garage did not run the model.
The repository is public under the MIT license and has source, tests, setup scripts, and build files. GitHub's release API reported version 0.1.39 on October 4 with three Windows archives and SHA-256 digests. Linux setup is available from the source tree.[1]
This host does not have the required consumer graphics card or enough free model storage for the documented route. We inspected the source tree, README, install guide, release record, and model card. We did not download the weights, install the engine, measure memory, or reproduce the speed table.
The Hacker News item was the discovery route. Its title and comments are not benchmark evidence.[4] Use the separator below to prepare a run card before a parameter badge turns into a purchasing decision.