Speculative Decoding Demo

Target: Qwen/Qwen2.5-3B-Instruct  ·  Draft: Qwen/Qwen2.5-0.5B-Instruct

Normal vs assisted generation. Automatically picks a device path where speculation can show a real wall-clock benefit.

32 256

Normal generation

Speculative (assisted) generation

Results

Run a comparison to see timing and speedup metrics.

Why speculative decoding helps (memory bandwidth)

Autoregressive decoding is usually memory-bandwidth bound: every new token streams the full target weights. FLOPs wait on memory.

Speculative decoding:

  1. A small draft proposes several tokens cheaply.
  2. The large target verifies the block in one forward pass.
  3. Accepted tokens are kept for roughly the cost of one target step.

Why this demo may run on CPU: on a fast GPU a 3B target is already quick, and a 0.5B draft is only ~6× smaller — draft overhead can erase the gains (speedup < 1). When that happens we re-run on CPU, where each target step is expensive and speculation’s fewer target passes show a clear wall-clock win. That is the same regime as large / offloaded models in production (the case HF highlights for assisted generation), without using accelerate offload hooks that break ZeroGPU.


Greedy decoding (do_sample=False). Uses generate(..., assistant_model=...). If a short GPU probe shows no win (typical for 3B+0.5B fully on GPU), the timed run switches to CPU so each target step is expensive and speculation can win.