IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Python class

KVCacheBuffer

KVCacheBuffer​

class max.nn.kv_cache.KVCacheBuffer(replicates_kv_across_tp, values, scales=None, values_per_layer=None, scales_per_layer=None, is_jenga=False)

source

Bases: KVCacheBufferInterface

A collection of KVCache buffers for one data-parallel replica.

Two buffer kinds are supported: values and (optionally, for FP8 quantization) scales. The length of each list corresponds to the tensor-parallel degree, with one buffer per TP shard.

replicates_kv_across_tp is True when the KV data is replicated identically across TP shards and False when it is sharded. The data is replicated in certain cases like TP + MLA, TP + MiniMaxM3IndexerAttn, etc.

Parameters:

all_buffers​

property all_buffers: list[Buffer]

source

Returns all value and scale buffers in a single flat list.

Returns:

A list containing every value buffer followed by every scale buffer (if scales are present).

is_jenga​

is_jenga: bool = False

source

Whether this buffer is associated with Jenga KV cache

TODO: Delete this field after reworking KVCacheBufferInterface.

replicates_kv_across_tp​

replicates_kv_across_tp: bool

source

scales​

scales: list[Buffer] | None = None

source

Per-TP-shard scale buffers for a quantized cache; None when unquantized.

scales_per_layer​

scales_per_layer: list[list[Buffer]] | None = None

source

Per-TP-shard, per-layer scale buffers for a quantized KV cache backed by per_layer_buffers (mirrors values_per_layer). scales[shard] aliases scales_per_layer[shard][0]. None for a single multi-layer scale buffer or an unquantized cache.

to_memory()​

to_memory()

source

Converts to offload-ready memory units, one per buffer kind.

Every buffer is re-viewed as 2-D uint8 pages so consumers can treat all caches uniformly regardless of dtype or shape.

Per-layer buffers are deliberately not enumerated – only each shard’s layer-0 alias – which is why allocate_buffers rejects per_layer_buffers alongside off-device connectors and DP > 1.

Returns:

One KVCacheMemory per kind (values, and scales if present).

Return type:

list[KVCacheMemory]

total_num_pages​

property total_num_pages: int

source

Returns the total number of pages across all values and scales.

values​

values: list[Buffer]

source

values_per_layer​

values_per_layer: list[list[Buffer]] | None = None

source

Per-TP-shard, per-layer value buffers when the pool uses per_layer_buffers.

values_per_layer[shard] is the list of single-layer buffers for that shard, and values[shard] aliases values_per_layer[shard][0] so the single-buffer values invariants (and consumers) stay valid. None for a normal single multi-layer buffer.