# How much extra VRAM speculative decoding needs in llama.cpp (MTP, DFlash, draft models)

Speculative decoding makes local LLMs noticeably faster, but it is not free: the draft (an MTP head, a DFlash drafter or a small draft model) needs its own VRAM. People regularly run out of memory after turning it on, and llama.cpp's automatic fitting does not always warn them. Here is where the extra memory goes, with exact numbers read from the GGUF files.

## What costs memory

For llama.cpp's `--spec-type draft-mtp`, `draft-dflash` or `draft-simple`, the extra VRAM is the sum of:

1.  **Draft weights.** The MTP block inside the main file is skipped at load time unless MTP is on (it is loaded with `TENSOR_SKIP` in `src/models/qwen35.cpp`), so it costs nothing when you don't use it. A separate draft file (`-md`) always costs its full size, minus `token_embd`, which llama.cpp keeps on the CPU.
    
2.  **The draft's KV cache.** The draft context gets the same `n_ctx` as the main model (`common/speculative.cpp`), and there is no flag to shrink it separately. Its cache type is set with `-ctkd` / `-ctvd`, independent of the main `-ctk` / `-ctv`.
    
3.  **Recurrent-state snapshots.** On hybrid models (Qwen3.x with Gated DeltaNet layers), MTP and DFlash keep `--spec-draft-n-max` extra copies of the recurrent state so rejected tokens can be rolled back. For Qwen3.8 27B one copy is 149.625 MiB.
    
4.  **A compute buffer** for the draft context, roughly 127–272 MiB in the CUDA logs people have posted, plus an F16 copy of the draft cache if you quantize it and use flash attention.
    

## Worked examples

**Qwen3.8 27B, unsloth UD-Q4\_K\_M, built-in MTP, 64K context, f16 caches,** `--spec-draft-n-max 3`**:**

| Part | Size |
| --- | --- |
| MTP block weights | 335 MB |
| Draft KV cache (1 layer × 4 KV heads × 256, at 64K) | 256 MB |
| 3 recurrent-state snapshots | 449 MB |
| Compute buffer (estimate) | 127–272 MB |
| **Extra VRAM** | **about 1.14–1.28 GB** |

The model alone needs about 20.2 GB at 64K, so plan on roughly 21.4 GB with MTP on.

**Same model with the z-lab DFlash2 Q4\_K\_M drafter:** about +1.67–1.81 GB. The drafter's weights are bigger (1.05 GiB), but its five layers use a 2,048-token sliding window, so its cache stays at about 50 MB whatever the context.

**Gemma 4 31B with Unsloth's F16 MTP assistant at 256K:** about +1.00–1.14 GB, almost all of it weights (896 MB). The Gemma 4 assistant reads the main model's KV cache instead of keeping its own, so its cache costs nothing and `-ctkd` / `-ctvd` have no effect on it.

## The trap: `-fit` may count the draft as zero

If you rely on llama.cpp's automatic fitting (no `-c`), the fitter measures the draft by creating a context for it on its own. Drafts that need the main context (the Gemma 4 assistant and DFlash files) fail that measurement, and `common/fit.cpp` logs:

```plaintext
failed to measure the memory of the extra model, fitting without it
```

It then sizes the main model's context without the draft model's own KV cache and context. The recurrent-state snapshots are still counted, because they are allocated inside the main model's context; they scale with `--spec-draft-n-max` and with `-np` (parallel slots), not with prompt length. Because the draft's own memory is missed, you can hit an out-of-memory error on the first long prompt ([llama.cpp #29521](https://github.com/ggml-org/llama.cpp/issues/29521) is a real example on a 128 GB Mac). The fix is to set the context yourself, for example `-c 262144 -np 1`, and leave room for the numbers above.

## A command that only uses flags that exist

```plaintext
llama-server -m Qwen3.8-27B-UD-Q4_K_M.gguf -ngl all -c 65536 -np 1 --spec-type draft-mtp --spec-draft-n-max 3
```

(There is no `-cd` for the draft context on current master; `-Cd` is the draft CPU mask.)

## Calculator

Disclosure: I built a free calculator that adds these parts up for Qwen3.8 27B and Gemma 4 31B with their MTP heads, DFlash drafters and small draft models, and prints the command: [modelvram.com/speculative-decoding-vram-calculator](https://modelvram.com/speculative-decoding-vram-calculator/). Byte counts come from the GGUF headers on Hugging Face, and every formula is checked against llama.cpp logs people have posted.

Updated 2026-09-29: clarified that -fit does count the recurrent-state snapshots and only misses the draft's own KV cache, thanks to a reader comment.
