Speculative decoding: the baseline

How a single worker decodes several tokens per forward pass: a cheap drafter guesses a block, the target model verifies the whole block at once, and the matching prefix is committed. (The distributed version is here: DSpec explainer.)
committed draft (cheap model) accepted by target bonus token (target sample) rejected ←/→ keys work too