For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python class
KVCacheInputsPerDevice
KVCacheInputsPerDevice
class max.nn.kv_cache.KVCacheInputsPerDevice(kv_blocks, cache_lengths, lookup_table, max_prompt_length, max_cache_length, kv_scales=None, scales_lookup_table=None, attention_dispatch_metadata=None, draft_attention_dispatch_metadata=None, mla_num_partitions=None, draft_mla_num_partitions=None, kv_blocks_per_layer=None, kv_scales_per_layer=None)
Bases: Generic[_Tensor, _Buffer]
Symbolic graph input types for a single device’s paged KV cache.
-
Parameters:
-
- kv_blocks (_Buffer)
- cache_lengths (_Tensor)
- lookup_table (_Tensor)
- max_prompt_length (_Tensor)
- max_cache_length (_Tensor)
- kv_scales (_Buffer | None)
- scales_lookup_table (_Tensor | None)
- attention_dispatch_metadata (_Tensor | None)
- draft_attention_dispatch_metadata (_Tensor | None)
- mla_num_partitions (_Tensor | None)
- draft_mla_num_partitions (_Tensor | None)
- kv_blocks_per_layer (list[_Buffer] | None)
- kv_scales_per_layer (list[_Buffer] | None)
attention_dispatch_metadata
attention_dispatch_metadata: _Tensor | None = None
The device’s attention dispatch metadata, as a rank-1 int64 tensor.
cache_lengths
cache_lengths: _Tensor
Per-request cache lengths, one rank-1 entry per request.
draft_attention_dispatch_metadata
draft_attention_dispatch_metadata: _Tensor | None = None
Attention dispatch metadata resolved for the draft-forward query width, as a rank-1 int64 tensor.
The draft_ prefix names the shorter query length, not the draft
model; a target cache leaf carries this variant too.
draft_mla_num_partitions
draft_mla_num_partitions: _Tensor | None = None
The analog of mla_num_partitions for the draft-forward query
width, carried by target cache leaves as well as draft ones.
flatten()
flatten()
Serialize fields into a flat list for graph input binding.
-
Return type:
-
list[_Tensor | _Buffer]
flatten_without_attention_dispatch_metadata()
flatten_without_attention_dispatch_metadata()
Serializes fields into a flat list, minus the attention dispatch metadata fields.
-
Return type:
-
list[_Tensor | _Buffer]
kv_blocks
kv_blocks: _Buffer
The device’s paged KV cache blocks.
kv_blocks_per_layer
One single-layer KV buffer per layer, used when the backing pool
allocates a standalone buffer per layer
(KVCacheParams.per_layer_buffers) instead of one multi-layer buffer.
kv_blocks aliases kv_blocks_per_layer[0] so single-buffer
consumers stay valid; a per-layer attention dispatch picks
kv_blocks_per_layer[layer_idx]. None (the default) for every
non-per-layer cache.
kv_scales
kv_scales: _Buffer | None = None
KV scales for FP8 quantization.
kv_scales_per_layer
One single-layer scale buffer per layer, the quantized-scale analog of
kv_blocks_per_layer (used with per_layer_buffers + a quantized
KV cache). kv_scales aliases kv_scales_per_layer[0]; a per-layer
attention dispatch picks kv_scales_per_layer[layer_idx]. None
for every non-per-layer / unquantized cache.
lookup_table
lookup_table: _Tensor
Per-request page lookup table, each row holding the block ids its request reads.
max_cache_length
max_cache_length: _Tensor
The batch’s maximum cache length, as a scalar tensor.
max_prompt_length
max_prompt_length: _Tensor
The batch’s maximum prompt length, as a scalar tensor.
mla_num_partitions
mla_num_partitions: _Tensor | None = None
Capturable-graph scalar the SM100 MLA dispatcher uses to align
grid-time partition decisions with the kernel’s divmod. Populated only
for MLA paths; None otherwise.
scales_lookup_table
scales_lookup_table: _Tensor | None = None
Page lookup table for kv_scales, present when the scales are paged
independently of the values so a request’s scale pages carry their own
ids. None means the two share one block-id space and lookup_table
resolves both, which is what every non-pooled cache does.
unflatten()
unflatten(it)
Reconstruct from a flat iterator produced by flatten.
Consumes next(it) in the same order flatten emits elements;
the two methods must stay in lock-step.
-
Parameters:
-
Return type: