A draft can come from repetition.
Prompt lookup decoding searches the prompt for an n-gram that matches the newest tokens. When it finds one, it proposes the tokens that followed the earlier match. The main model checks those candidates through speculative decoding, so accepted tokens still come from the target model's distribution.[1]
The method fits input-grounded work because names, phrases, and code often move from the prompt into the answer. The original project reported about 2.4 times higher throughput on its summarization and context-question-answering tests. It used greedy decoding, Mistral 7B, one A100, and 100 examples per task.[2]
Prompt lookup is a reuse lane. It helps when the answer repeats the input. It has less to grab when the model must invent every next phrase.
The lookup table is part of the hot path.
llama.cpp keeps context, dynamic, and static n-gram caches. Its example documents the controls for minimum and maximum n-gram size and the number of draft tokens.[1] Hayder Tirmazi profiled the cache path and found avoidable map copies, cache-unfriendly containers, and work that could be rejected before scoring every candidate.[3]
The reported gains are specific. On an Apple M4 Pro, with WikiText-103 replayed through llama-lookup-stats, the changes made drafting per drafted token faster, cut static-cache load time, and reduced peak memory. Daniel Lemire's later threshold check made that drafting path faster again in the same benchmark. The article reports medians of three runs and publishes the scripts and raw tables.[4]
Forty-two times is not forty-two times more tokens.
The headline measures time inside the draft lookup path. The expensive target-model verification still runs. End-to-end gain depends on prompt overlap, acceptance rate, batch shape, model cost, cache size, and hardware. A tiny draft subsystem can become much faster without moving the full request by the same ratio.
The implementation work also remains on open pull requests in a fork as of September 27. It is useful evidence, not an upstream release note. Benchmark the exact revision you plan to run and keep the baseline beside it.[3]
Use four clocks.
Record overlap and accepted draft tokens first. Then time draft work, target-model verification, and the complete request separately. Add peak memory and cache-load time when a static corpus is involved. If the request gets faster, the split clocks explain why. If it does not, they show whether lookup cost, low acceptance, or model verification ate the gain.
The gearbox below creates a test plan. Closing every clamp means the plan asks for each measurement. It does not mean a model, cache, or workload passed.