Pimp My IDE / garage dispatch
Back to garage
September 30, 2026 | local inference / benchmarks / agent workloads

Your fastest kernel can still lose the coding shift.

Magnitude tunes inference kernels on the machine that runs them. That is a useful change. It also makes the warm-up, workload, and retained tuning state part of the benchmark.

Pin the exact model and engine. Separate cold setup from warm speed. Then replay the agent task that has to survive both.

Hardware-specific tuning changes the test contract.

Magnitude describes an open-source inference engine that compiles and tunes kernels on the local device before a model runs. Its launch post reports matched results against llama.cpp for one Qwen 3.6 35B A3B four-bit setup at 64,000 tokens of context. The reported gains differ by machine and phase. The Metal decode result is much larger than the CUDA decode result.[1]

Those are vendor measurements for two named systems. They do not establish the result for another GPU, model, quantization, context length, or agent loop. The tuning mechanism makes that boundary more important because the generated kernel configuration belongs to one device and software stack.

A tuned engine is a machine plus the tuning record, not a portable speed adjective.

Cold setup and warm inference answer different questions.

A developer feels model download, load time, kernel tuning, prompt ingestion, token generation, tool pauses, and unload behavior. A steady-state token rate measures only part of that path. Magnitude's 0.2.2 release makes the distinction concrete. The app stopped running GPU kernels during its model assessment because that assessment could hang or fail on some hardware. It now estimates model speed from memory bandwidth.[2]

That estimate can help choose a model. It is not the measured result of a coding task. Record the estimate as discovery data. Record cold startup, first response, warm response, and task completion as separate observations.

Prefill and decode are separate lanes.

llama-bench names prompt processing, text generation, and the combined path as different tests. It repeats tests, reports average tokens per second and standard deviation, and runs a warm-up unless the operator disables it. Its documentation also says the measurements omit tokenization and sampling time.[3]

That makes it useful for engine work and incomplete for an agent decision. A coding agent may carry a large system prompt, reuse a prefix, call tools, wait on a repository, and resume several times. Measure engine phases. Then measure the task around them.

Use the workload that will own the machine.

vLLM's benchmark guide separates online service tests, offline throughput, prefix caching, trace replay, and other workloads. The guide says its included benchmarks mainly test vLLM functions and regressions. It points production-server work to a separate benchmarking framework.[4]

The lesson is not that one framework is correct. The lesson is to name the question. A single interactive agent needs first-response latency and stable warm turns. A long-context review needs prefill, cache behavior, and memory headroom. Concurrent agents need tail latency and proof that the desktop remains usable.

Interactive makeover / workload heat-soak bay

Build the benchmark around the shift

Traditional purpose replaced: one tokens-per-second screenshot. Better version: a native workload selector and four ordered test fields drive one physical heat-soak dial, a connected evidence rail, and a copyable run card.

Choose the load case

The selector changes the required measurements. The test fields choose what the run card will request. They do not report completed evidence.

Agent workload
Run-card fields
4 run-card fields openSingle-agent loop
Heat-soak circuit / requested evidence

Local inference heat-soak bay

The load case is selected. The artifact field is open.

Name the engine revision, model hash, quantization, backend, device, driver, and launch settings.

All four fields means the run-card structure is ready. It does not mean a kernel was tuned, a benchmark ran, a task succeeded, or one engine beat another.

One result / four records

Keep the machine attached to the number.

01 / ARTIFACT

Exact bits

Record the engine revision, model file hash, quantization, backend, driver, device, and launch settings.

02 / SOAK

Cold and warm

Separate discovery, tuning, model load, first response, and settled repeats. State what persisted between runs.

03 / LOAD

Agent-shaped work

Name prompt size, context depth, output length, tool turns, session count, cache policy, and repetitions.

04 / REPLAY

Task result

Keep task success, first-response latency, total time, memory, failures, variance, and retained output together.

Sources read

Source log and evidence boundary
  1. Magnitude repository and README, read September 30, 2026. The project describes local kernel compilation and tuning, dynamic memory allocation, shared prefix caches, supported systems, and vendor benchmark results for one Metal machine and one CUDA machine.
  2. Launch HN discussion for Magnitude, item 49911995, posted and read September 30, 2026. The launch text supplies the named model, quantization, context length, disabled speculative decoding, hardware, prefill rates, decode rates, and per-agent memory claims. We verified the discussion ID through the Hacker News API.
  3. Magnitude CLI 0.2.2 release, published October 1 and read September 30, 2026 in Eastern time. The release says model assessment now estimates speed from device memory bandwidth because running kernels during assessment could hang or fail on some hardware. It also lists agent-protocol and prompt-cache fixes.
  4. llama.cpp llama-bench documentation, read September 30, 2026. It documents prompt-processing, text-generation, combined tests, warm-up behavior, repetitions, averages, standard deviation, context depth, and the omission of tokenization and sampling time.
  5. vLLM Benchmark CLI documentation, read September 30, 2026. It separates online, offline, prefix-cache, trace-replay, and other test modes. It also states that the included benchmarks mainly evaluate vLLM functions and regressions, and points production-server testing elsewhere.

Evidence boundary: Magnitude's performance figures come from its authors. We inspected the repository, release, launch text, llama.cpp benchmark contract, and vLLM benchmark guide. We did not install the release or reproduce its measurements because this host has no designated test GPU, pinned model file, or matched baseline setup. The interactive bay writes a test plan. It reports no measured speed.