Pimp My IDE / inference inspection
Back to garage
October 4, 2026 | Local models / memory placement / benchmarks

Split the badge from the machine.

A 125B label, a 6B active path, and 100 tokens per second describe different parts of an inference run. Put them on separate instruments.

Parameter count names the model. Placement says where its bytes wait. Prompt processing and text generation measure different work.

GPUHot experts and active computeFast lane
RAMResident expert weights and host workCapacity lane
SSDModel files and offloaded tablesStorage lane

The 125B badge is real. It is not a memory map.

Qwen lists Qwen3.8-Flash-Next as a 125-billion-parameter language model with 6 billion parameters activated for each token. The model card also names 51 billion n-gram embedding parameters, a 4-billion-parameter multi-token prediction layer, 512 experts, and 11 active experts per mixture layer.[2]

Those numbers answer different questions. Total parameters describe stored model capacity. Activated parameters describe one compute path. Neither number says how many bytes sit in graphics memory, system memory, or storage after quantization.

Active parameters explain work per token. They do not erase the weights that still need a home.

Strata turns placement into the product.

Strata's public design keeps frequently used experts on the graphics card. It keeps the full expert set in system memory and puts a large lookup table on storage. The project says its normal route needs at least 12 GB of graphics memory, 32 GB of system memory, and about 80 GB of free disk. It says setup downloads about 70 GB and loads 35 to 55 GB into system memory.[1]

The project publishes measured prompt-processing and generation rates for named quantizations on an RTX 5070 and an RX 9070 XT. It also records the engine versions, a 32K-token prompt, and a 4K-token answer. Those details make the table useful. The headline alone does not preserve them.[1]

Prompt speed and answer speed need separate dials.

llama.cpp's benchmark tool defines prompt processing, text generation, and a combined test as separate jobs. It repeats tests, reports an average and standard deviation, accepts a context depth, and records placement controls such as GPU layers and CPU mixture-of-experts layers. Its documentation also says benchmark timing excludes tokenization and sampling.[3]

A local coding run needs one more timing boundary. Record model load, first response, prompt processing, text generation, and wall-clock task completion. A fast decode cannot repay a long model load if the process restarts for every task. A fast prompt pass cannot prove that edits, tool calls, or tests finish sooner.

Quantization is part of the model identity.

Strata offers several compressed variants. Its README says smaller variants run faster while larger variants keep more quality. The Coder option removes half of the experts to fit 32 GB of system memory. The project also warns that a larger 4-bit option can read weights from storage when system memory is tight, which cuts generation speed.[1]

Write the exact model family, quantization, engine revision, draft model, context size, cache format, and placement beside every result. Without that row, two runs with the same 125B badge may test different artifacts and different routes through the machine.

The artifact exists. This garage did not run the model.

The repository is public under the MIT license and has source, tests, setup scripts, and build files. GitHub's release API reported version 0.1.39 on October 4 with three Windows archives and SHA-256 digests. Linux setup is available from the source tree.[1]

This host does not have the required consumer graphics card or enough free model storage for the documented route. We inspected the source tree, README, install guide, release record, and model card. We did not download the weights, install the engine, measure memory, or reproduce the speed table.

The Hacker News item was the discovery route. Its title and comments are not benchmark evidence.[4] Use the separator below to prepare a run card before a parameter badge turns into a purchasing decision.

Interactive makeover / parameter badge separator

Build the load sheet.

This replaces one large model badge with four native review controls, a connected evidence bus, and a copyable run card. It prepares the test. It does not measure your machine.

Claim plates

Plan builder

Select each record that your planned run will capture. A selected plate marks the structure. It does not mean the value has been measured.

Load-sheet records

Evidence manifold

Structure only
IDENTITYPLACEMENTWORKLOADTIMING
LOAD SHEET
0 / 4No record is selected.
Load sheet open

No record is selected. Measurements and artifacts are still required.

This component drafts a load sheet. It does not inspect hardware, install software, download weights, run inference, judge model quality, or verify a throughput claim.

Sources read

Source log and evidence boundary
  1. Strata repository and README, read October 4, 2026. The project documents its memory-placement design, hardware requirements, model variants, install path, API routes, and project-measured prompt and generation speeds. We also read its install guide and inspected the v0.1.39 release. GitHub's API supplied the release date, archive sizes, and digests.
  2. Qwen3.8-Flash-Next model card, read October 4, 2026. Qwen supplies the parameter, activation, expert, architecture, context, license, and vendor benchmark claims. Those claims do not verify Strata's implementation or performance.
  3. llama.cpp llama-bench documentation, read October 4, 2026. It defines prompt-processing, text-generation, and combined tests. It also documents repetitions, context depth, placement controls, output fields, and timing exclusions.
  4. Hacker News item 49953495, resolved through the official API and read October 4, 2026. It was the discovery route. Its score, title, and comments are not used as independent product evidence.

Evidence boundary. Qwen documents the model. Strata documents its engine and publishes source and release artifacts. llama.cpp documents its benchmark tool. Pimp My IDE designed the four-record load sheet. We did not run the model, reproduce a throughput result, test output quality, or inspect installed network behavior.