For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python class
SequentialDriver
SequentialDriver
class max.pipelines.speculative.driver.SequentialDriver(target, proposer, *, target_model, draft_model, devices, data_parallel_degree, input_spec, speculative_config=None, enable_structured_output=False)
Bases: Module
Drives the sequential speculative-decoding loop.
Each iteration runs merge, verify, mask, accept, shift, propose, and pack.
-
Parameters:
-
- target (SpecDecodeTarget[SequentialBatch])
- proposer (SequentialProposer)
- target_model (Module)
- draft_model (Module)
- devices (Sequence[DeviceRef])
- data_parallel_degree (int)
- input_spec (SpecDecodeInputTypeSpec)
- speculative_config (SpeculativeConfig | None)
- enable_structured_output (bool)
input_types()
input_types(kv_params)
Builds the unified spec-decode graph signature.
See build_spec_decode_input_types for the canonical ordering.
-
Parameters:
-
kv_params (KVCacheParamInterface)
-
Return type:
-
tuple[TensorType | BufferType, …]