Hardware-specific tuning changes the test contract.
Magnitude describes an open-source inference engine that compiles and tunes kernels on the local device before a model runs. Its launch post reports matched results against llama.cpp for one Qwen 3.6 35B A3B four-bit setup at 64,000 tokens of context. The reported gains differ by machine and phase. The Metal decode result is much larger than the CUDA decode result.[1]
Those are vendor measurements for two named systems. They do not establish the result for another GPU, model, quantization, context length, or agent loop. The tuning mechanism makes that boundary more important because the generated kernel configuration belongs to one device and software stack.
A tuned engine is a machine plus the tuning record, not a portable speed adjective.
Cold setup and warm inference answer different questions.
A developer feels model download, load time, kernel tuning, prompt ingestion, token generation, tool pauses, and unload behavior. A steady-state token rate measures only part of that path. Magnitude's 0.2.2 release makes the distinction concrete. The app stopped running GPU kernels during its model assessment because that assessment could hang or fail on some hardware. It now estimates model speed from memory bandwidth.[2]
That estimate can help choose a model. It is not the measured result of a coding task. Record the estimate as discovery data. Record cold startup, first response, warm response, and task completion as separate observations.
Prefill and decode are separate lanes.
llama-bench names prompt processing, text generation, and the combined path as different tests. It repeats tests, reports average tokens per second and standard deviation, and runs a warm-up unless the operator disables it. Its documentation also says the measurements omit tokenization and sampling time.[3]
That makes it useful for engine work and incomplete for an agent decision. A coding agent may carry a large system prompt, reuse a prefix, call tools, wait on a repository, and resume several times. Measure engine phases. Then measure the task around them.
Use the workload that will own the machine.
vLLM's benchmark guide separates online service tests, offline throughput, prefix caching, trace replay, and other workloads. The guide says its included benchmarks mainly test vLLM functions and regressions. It points production-server work to a separate benchmarking framework.[4]
The lesson is not that one framework is correct. The lesson is to name the question. A single interactive agent needs first-response latency and stable warm turns. A long-context review needs prefill, cache behavior, and memory headroom. Concurrent agents need tail latency and proof that the desktop remains usable.