IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Python module

max.nn.kv_cache

Cache configuration​

KVCacheBufferA collection of KVCache buffers for one data-parallel replica.
KVCacheParamInterfaceInterface for KV cache parameters.
KVCacheParamsConfiguration parameters for key-value cache management in transformer models.
MHAKVCacheParamsKV cache parameters for multi-head attention (MHA).
MLAKVCacheParamsKV cache parameters for multi-latent attention (MLA).
MSAKVCacheParamsKV cache parameters for multi-step attention (MSA).
KVCacheQuantizationConfigConfiguration for KVCache quantization.
KVConnectorTypeIdentifies which off-device backing store the KV cache uses.
KVCacheMemoryOne logical (child, kind) KV tensor as per-TP-shard uint8 views.
MultiKVCacheParamsAggregates multiple KV cache parameter sets into a recursive tree.
PagedKVLeafRegionA leaf the graph reaches through a per-forward page table.

Cache inputs​

KVCacheInputsSymbolic graph input types for a leaf KV cache.
KVCacheInputsPerDeviceSymbolic graph input types for a single device's paged KV cache.
BatchCharacteristicsUpper-bound batch shape used to prepare decode attention metadata.
PagedCacheValuesalias of KVCacheInputsPerDevice[TensorValue, BufferValue]

Attention dispatch​

AttnKeyA resolved decode-attention dispatch shape.
AttnKeyInterfaceCommon base for resolved attention keys.
MHAAttnKeyDecode dispatch metadata for multi-head attention (MHA).
MLAAttnKeyDecode dispatch metadata for multi-latent attention (MLA).
MSAAttnKeyDecode dispatch metadata for multi-step attention (MSA).

Metrics​

KVCacheMetricsMetrics for the KV cache.

Functions​

build_max_lengths_tensorsBuilds two [1] uint32 scalar buffers of maximum lengths.
compute_max_seq_len_fitting_in_cacheComputes the maximum sequence length that can fit in the available memory.
compute_num_device_blocksComputes the number of blocks that can be allocated based on the available cache memory.
estimated_memory_sizeComputes the estimated memory size of the KV cache used by all replicas.
padded_lut_colsRounds a page lookup-table inner dim up to a kernel-safe width.
spec_decode_cache_slackComputes the extra KV positions a request may occupy past max_seq_len.