Pimp My IDE / Garage dispatch
Back to garage
September 27, 2026 | local inference / prompt lookup / benchmark discipline

The cheapest draft model may be the prompt you already have.

Prompt lookup proposes repeated phrases from the input, then lets the main model verify them in one pass. It can help code edits, extraction, and document work. It offers little benefit when open-ended output does not repeat the prompt.

The take. Route prompt lookup by workload. Measure accepted drafts and end-to-end tokens per second. Do not turn a faster cache lookup into a blanket claim about model speed.

A draft can come from repetition.

Prompt lookup decoding searches the prompt for an n-gram that matches the newest tokens. When it finds one, it proposes the tokens that followed the earlier match. The main model checks those candidates through speculative decoding, so accepted tokens still come from the target model's distribution.[1]

The method fits input-grounded work because names, phrases, and code often move from the prompt into the answer. The original project reported about 2.4 times higher throughput on its summarization and context-question-answering tests. It used greedy decoding, Mistral 7B, one A100, and 100 examples per task.[2]

Prompt lookup is a reuse lane. It helps when the answer repeats the input. It has less to grab when the model must invent every next phrase.

The lookup table is part of the hot path.

llama.cpp keeps context, dynamic, and static n-gram caches. Its example documents the controls for minimum and maximum n-gram size and the number of draft tokens.[1] Hayder Tirmazi profiled the cache path and found avoidable map copies, cache-unfriendly containers, and work that could be rejected before scoring every candidate.[3]

The reported gains are specific. On an Apple M4 Pro, with WikiText-103 replayed through llama-lookup-stats, the changes made drafting per drafted token faster, cut static-cache load time, and reduced peak memory. Daniel Lemire's later threshold check made that drafting path faster again in the same benchmark. The article reports medians of three runs and publishes the scripts and raw tables.[4]

Forty-two times is not forty-two times more tokens.

The headline measures time inside the draft lookup path. The expensive target-model verification still runs. End-to-end gain depends on prompt overlap, acceptance rate, batch shape, model cost, cache size, and hardware. A tiny draft subsystem can become much faster without moving the full request by the same ratio.

The implementation work also remains on open pull requests in a fork as of September 27. It is useful evidence, not an upstream release note. Benchmark the exact revision you plan to run and keep the baseline beside it.[3]

Use four clocks.

Record overlap and accepted draft tokens first. Then time draft work, target-model verification, and the complete request separately. Add peak memory and cache-load time when a static corpus is involved. If the request gets faster, the split clocks explain why. If it does not, they show whether lookup cost, low acceptance, or model verification ate the gain.

The gearbox below creates a test plan. Closing every clamp means the plan asks for each measurement. It does not mean a model, cache, or workload passed.

Interactive makeover / benchmark routing

Prompt lookup gearbox

Traditional purpose replaced: enable prompt lookup and quote one speedup. Better version: select the workload, close four measurement clamps, and copy a test card that keeps subsystem timing separate from the full request.

Set the reuse lane

Use one workload gear. The four square clamps select independent measurements.

Workload gear
0/4Route broken at Overlap. No benchmark evidence recorded.

Teaching display. The count shows selected test sections, not measured speed or readiness.

Benchmark template0 SECTIONS SELECTED

Copy the lookup test card

Selection adds required fields. It does not fill them or prove a speedup.

What completion means. Four measurement sections are present. Model details, prompts, commands, timings, acceptance, memory, and outputs still need real values.

Sources and limits

Open the source log
  1. llama.cpp prompt lookup example, read September 27, 2026. It documents the example and its n-gram and draft-count controls.
  2. Apoorv Saxena's Prompt Lookup Decoding repository, read September 27, 2026. It explains the method and reports task-specific tests on an A100 with Mistral 7B. Those results do not predict another workload or machine.
  3. Hayder Tirmazi, "42x Faster Prompt Lookup Drafting in llama.cpp", published September 26 and read September 27, 2026. It explains the cache changes, test machine, corpus, metrics, and limits. Its linked implementation pull requests were open in a fork when checked.
  4. ngram-cache-bench repository, read September 27, 2026. It pins the benchmark variants, corpus steps, settings, raw run tables, and plot generation.
  5. Hacker News discussion 49859982, checked through the official HN API on September 27, 2026. It surfaced the article. Votes and comments are attention signals, not benchmark evidence.

Source boundary. Method claims come from the original project and llama.cpp documentation. Performance numbers come from the named authors on their stated hardware and datasets. The benchmark card is editorial guidance. This page did not run a model benchmark.