IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Python module

max.pipelines.speculative.driver

Model-agnostic driver for sequential (eagle / MTP) speculative decoding.

Every sequential unified spec-decode module in the tree runs the same seven phases โ€“ merge, verify, mask, accept, shift, propose, pack โ€“ and differs only in how it calls its target, how it calls its draft, and what per-step cache bookkeeping the draft needs. SequentialDriver owns the phases; a model contributes a SpecDecodeTarget and a SequentialProposer.

The two adapters are deliberately not Module subclasses. The driver registers the target and draft modules under the same attribute names the hand-written modules used, so weight loading, state_dict keys and the weights registry are unchanged.

Sequential driverโ€‹

DecodeKVSwapA single per-step edit to the draft's PagedCacheValues.
DraftStepInputLoop-carried state between draft steps.
ProposedOne draft invocation's contribution.
SequentialBatchOne spec-decode iteration's graph inputs, merged and broadcast.
SequentialDriverDrives the sequential speculative-decoding loop.
SequentialProposerA draft that emits one token per step, K steps deep.