For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python class
MLAAttnKey
MLAAttnKey
class max.nn.kv_cache.MLAAttnKey(batch_size, max_prompt_length, num_partitions)
Bases: AttnKey
Decode dispatch metadata for multi-latent attention (MLA).
pack_into_buffer()
pack_into_buffer(device, max_cache_valid_length)
Returns the accelerator dispatch buffer MLA decode kernels read, holding batch size, prompt width, and partition count.