IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Python class

RecurrentKVLeafRegion

RecurrentKVLeafRegion​

class max.nn.kv_cache.RecurrentKVLeafRegion(leaf_id, group_id, bytes_per_page, region, *, row_bytes=1)

source

Bases: KVLeafRegion

A leaf addressed by row: one block holds one request’s whole state.

Parameters:

blocks_to_reserve()​

blocks_to_reserve(num_blocks, *, enable_prefix_caching)

source

Returns the live block, and a checkpoint behind it with caching on.

Without prefix caching a state never checkpoints.

Parameters:

  • num_blocks (int)
  • enable_prefix_caching (bool)

Return type:

int

bound_row_copies()​

bound_row_copies(src, dst)

source

Returns the state rows to copy, keyed where the pool is bound.

Parameters:

Return type:

Mapping[str, tuple[range, range]]

bound_row_span()​

bound_row_span(block)

source

Returns the rows one block’s layers occupy.

Parameters:

block (int)

Return type:

Mapping[str, range]

region​

region: RecurrentStateRegion

source

The state region that names this leaf and sizes its rows.

staged_input_shapes()​

staged_input_shapes(batch_size, num_blocks)

source

Returns the row table each layer slices.

The table is layer-major, one layer per row, so a layer’s slice is a contiguous row and the graph keeps it a view instead of materializing a gather kernel. num_blocks is unused: a state is addressed by row, not by block.

Parameters:

  • batch_size (int)
  • num_blocks (int)

Return type:

Mapping[str, tuple[tuple[int, …], DType]]

write_staged_inputs()​

write_staged_inputs(plans, into)

source

Folds each request’s block into the rows its layers index.

Parameters:

Return type:

None