For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python class
DecodeKVSwap
DecodeKVSwap
class max.pipelines.speculative.driver.DecodeKVSwap(value, names=<not given>, *values, module=None, qualname=None, type=None, start=1, boundary=None)
Bases: Enum
A single per-step edit to the draft’s PagedCacheValues.
The draft’s step-0 pass runs over the whole corrected sequence, so its
query length is whatever the merged batch holds; steps 1..K-1 run one
token per request, q = 1. PagedCacheValues carries a dispatch
buffer for each, the q = 1 one under a draft_ prefix – which
names the shorter query length, not the draft model; the target’s cache
leaf carries a draft_ variant too. Selecting between them is a
declaration rather than an inline replace per model.
DRAFT_ATTENTION_DISPATCH_METADATA
DRAFT_ATTENTION_DISPATCH_METADATA = 'draft_attention_dispatch_metadata'
DRAFT_MLA_NUM_PARTITIONS
DRAFT_MLA_NUM_PARTITIONS = 'draft_mla_num_partitions'
MAX_PROMPT_LENGTH_ONE
MAX_PROMPT_LENGTH_ONE = 'max_prompt_length_one'
ZERO_CACHE_LENGTHS
ZERO_CACHE_LENGTHS = 'zero_cache_lengths'
no past to attend over.
Applied by the driver for DraftCache.PER_DEPTH rather than
declared, since it only ever means that.
-
Type:
-
Start each step’s window at zero