Speculative Decoding Demo
Target: Qwen/Qwen2.5-3B-Instruct · Draft: Qwen/Qwen2.5-0.5B-Instruct
Normal vs assisted generation. Automatically picks a device path where speculation can show a real wall-clock benefit.
Normal generation
Speculative (assisted) generation
Results
Run a comparison to see timing and speedup metrics.
Why speculative decoding helps (memory bandwidth)
Autoregressive decoding is usually memory-bandwidth bound: every new token streams the full target weights. FLOPs wait on memory.
Speculative decoding:
- A small draft proposes several tokens cheaply.
- The large target verifies the block in one forward pass.
- Accepted tokens are kept for roughly the cost of one target step.
Why this demo may run on CPU: on a fast GPU a 3B target is already quick, and a 0.5B draft is only ~6× smaller — draft overhead can erase the gains (speedup < 1). When that happens we re-run on CPU, where each target step is expensive and speculation’s fewer target passes show a clear wall-clock win. That is the same regime as large / offloaded models in production (the case HF highlights for assisted generation), without using accelerate offload hooks that break ZeroGPU.
Greedy decoding (do_sample=False). Uses generate(..., assistant_model=...).
If a short GPU probe shows no win (typical for 3B+0.5B fully on GPU), the timed
run switches to CPU so each target step is expensive and speculation can win.