IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Python class

KVCacheInputsPerDevice

KVCacheInputsPerDevice​

class max.nn.kv_cache.KVCacheInputsPerDevice(kv_blocks, cache_lengths, lookup_table, max_prompt_length, max_cache_length, kv_scales=None, scales_lookup_table=None, attention_dispatch_metadata=None, draft_attention_dispatch_metadata=None, mla_num_partitions=None, draft_mla_num_partitions=None, kv_blocks_per_layer=None, kv_scales_per_layer=None)

source

Bases: Generic[_Tensor, _Buffer]

Symbolic graph input types for a single device’s paged KV cache.

Parameters:

  • kv_blocks (_Buffer)
  • cache_lengths (_Tensor)
  • lookup_table (_Tensor)
  • max_prompt_length (_Tensor)
  • max_cache_length (_Tensor)
  • kv_scales (_Buffer | None)
  • scales_lookup_table (_Tensor | None)
  • attention_dispatch_metadata (_Tensor | None)
  • draft_attention_dispatch_metadata (_Tensor | None)
  • mla_num_partitions (_Tensor | None)
  • draft_mla_num_partitions (_Tensor | None)
  • kv_blocks_per_layer (list[_Buffer] | None)
  • kv_scales_per_layer (list[_Buffer] | None)

attention_dispatch_metadata​

attention_dispatch_metadata: _Tensor | None = None

source

The device’s attention dispatch metadata, as a rank-1 int64 tensor.

cache_lengths​

cache_lengths: _Tensor

source

Per-request cache lengths, one rank-1 entry per request.

draft_attention_dispatch_metadata​

draft_attention_dispatch_metadata: _Tensor | None = None

source

Attention dispatch metadata resolved for the draft-forward query width, as a rank-1 int64 tensor.

The draft_ prefix names the shorter query length, not the draft model; a target cache leaf carries this variant too.

draft_mla_num_partitions​

draft_mla_num_partitions: _Tensor | None = None

source

The analog of mla_num_partitions for the draft-forward query width, carried by target cache leaves as well as draft ones.

flatten()​

flatten()

source

Serialize fields into a flat list for graph input binding.

Return type:

list[_Tensor | _Buffer]

flatten_without_attention_dispatch_metadata()​

flatten_without_attention_dispatch_metadata()

source

Serializes fields into a flat list, minus the attention dispatch metadata fields.

Return type:

list[_Tensor | _Buffer]

kv_blocks​

kv_blocks: _Buffer

source

The device’s paged KV cache blocks.

kv_blocks_per_layer​

kv_blocks_per_layer: list[_Buffer] | None = None

source

One single-layer KV buffer per layer, used when the backing pool allocates a standalone buffer per layer (KVCacheParams.per_layer_buffers) instead of one multi-layer buffer. kv_blocks aliases kv_blocks_per_layer[0] so single-buffer consumers stay valid; a per-layer attention dispatch picks kv_blocks_per_layer[layer_idx]. None (the default) for every non-per-layer cache.

kv_scales​

kv_scales: _Buffer | None = None

source

KV scales for FP8 quantization.

kv_scales_per_layer​

kv_scales_per_layer: list[_Buffer] | None = None

source

One single-layer scale buffer per layer, the quantized-scale analog of kv_blocks_per_layer (used with per_layer_buffers + a quantized KV cache). kv_scales aliases kv_scales_per_layer[0]; a per-layer attention dispatch picks kv_scales_per_layer[layer_idx]. None for every non-per-layer / unquantized cache.

lookup_table​

lookup_table: _Tensor

source

Per-request page lookup table, each row holding the block ids its request reads.

max_cache_length​

max_cache_length: _Tensor

source

The batch’s maximum cache length, as a scalar tensor.

max_prompt_length​

max_prompt_length: _Tensor

source

The batch’s maximum prompt length, as a scalar tensor.

mla_num_partitions​

mla_num_partitions: _Tensor | None = None

source

Capturable-graph scalar the SM100 MLA dispatcher uses to align grid-time partition decisions with the kernel’s divmod. Populated only for MLA paths; None otherwise.

scales_lookup_table​

scales_lookup_table: _Tensor | None = None

source

Page lookup table for kv_scales, present when the scales are paged independently of the values so a request’s scale pages carry their own ids. None means the two share one block-id space and lookup_table resolves both, which is what every non-pooled cache does.

unflatten()​

unflatten(it)

source

Reconstruct from a flat iterator produced by flatten.

Consumes next(it) in the same order flatten emits elements; the two methods must stay in lock-step.

Parameters:

it (Iterator[Any])

Return type:

KVCacheInputsPerDevice[TensorValue, BufferValue]