For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python module
max.pipelines.speculative.driver
Model-agnostic driver for sequential (eagle / MTP) speculative decoding.
Every sequential unified spec-decode module in the tree runs the same seven
phases โ merge, verify, mask, accept, shift, propose, pack โ and differs only
in how it calls its target, how it calls its draft, and what per-step cache
bookkeeping the draft needs. SequentialDriver owns the phases; a model
contributes a SpecDecodeTarget and a SequentialProposer.
The two adapters are deliberately not Module subclasses. The driver
registers the target and draft modules under the same attribute names the
hand-written modules used, so weight loading, state_dict keys and the
weights registry are unchanged.
Sequential driverโ
DecodeKVSwap | A single per-step edit to the draft's PagedCacheValues. |
|---|---|
DraftStepInput | Loop-carried state between draft steps. |
Proposed | One draft invocation's contribution. |
SequentialBatch | One spec-decode iteration's graph inputs, merged and broadcast. |
SequentialDriver | Drives the sequential speculative-decoding loop. |
SequentialProposer | A draft that emits one token per step, K steps deep. |