IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Python class

RecurrentKVGroupCoordinator

RecurrentKVGroupCoordinator​

class max.pipelines.kv_cache.paged_kv_cache.recurrent_coordinator.RecurrentKVGroupCoordinator(pools, leaf_id, group_id, page_size=0, enable_prefix_caching=False, *, rows=<factory>, successors=<factory>, loading=<factory>)

source

Bases: KVGroupCoordinatorInterface

A group whose entry is one state, held in the last slot of a row.

row[-1] is the block the recurrence runs in, and it carries no hash. Every block behind it is null, a published checkpoint, or one a connector hit is still loading. The recurrence reads and writes one block, so reaching a boundary publishes the block it ran in and copies it into a successor rather than snapshotting it.

Parameters:

advance()​

advance(req_id, num_committed_blocks, replica_idx)

source

Frees every published block behind the live one, nulling its slot.

Parameters:

Return type:

None

blocks_held_of_connector_hit()​

blocks_held_of_connector_hit(num_hit_blocks)

source

Returns one block, the state at the hit’s depth.

Parameters:

num_hit_blocks (int)

Return type:

int

blocks_to_allocate()​

blocks_to_allocate(req_id, num_required_blocks)

source

Returns the live block the row lacks, and its missing successor.

Parameters:

Return type:

dict[str, int]

checkpoint()​

checkpoint(ctx, replica_idx)

source

Publishes the block just run in and fills the one that succeeds it.

The block the forward ran in already holds the state the boundary hash names, so it is published where that hash lands rather than snapshotted. The request continues in a freshly drawn block, copied from it. Runs before the commit, so the copy reads a block nothing has freed yet.

Empty unless the forward ended exactly on a block boundary and the row holds no unpublished predecessor already.

Parameters:

Return type:

Mapping[str, tuple[int | None, int]]

claim_hit_blocks()​

claim_hit_blocks(desired_hashes, replica_idx)

source

Takes the block at the granted block count, nulling the slots behind it.

Returns empty rows when no published state stands there.

Parameters:

Return type:

dict[str, list[LittleKVCacheBlock]]

claimable_hashes()​

claimable_hashes(desired_hashes)

source

Returns the deepest hash: a state is one boundary, not a run.

Parameters:

desired_hashes (Sequence[bytes])

Return type:

Sequence[bytes]

commit()​

commit(req_id, hashes, last_block, replica_idx)

source

Publishes the checkpoints below last_block.

Forgets a loading checkpoint that a twin holding its hash replaced.

Parameters:

Return type:

None

enable_prefix_caching​

enable_prefix_caching: bool = False

source

Whether the state checkpoints at page boundaries.

Without prefix caching nothing is published, so the state stays in the block it ran in.

extend()​

extend(req_id, hit_blocks, loaded_blocks, replica_idx)

source

Appends the hit, recording a loaded checkpoint as loading.

Parameters:

Return type:

None

forward_blocks()​

forward_blocks(batch, num_blocks)

source

Returns the block each request’s recurrence runs in.

Parameters:

Return type:

dict[str, list[list[int]]]

grow()​

grow(req_id, num_required_blocks, replica_idx)

source

Draws the live block and successor the request lacks.

Also pads the row out.

Parameters:

Return type:

None

live_blocks()​

live_blocks(req_id)

source

Returns the block per leaf the recurrence runs in, if one is drawn.

Parameters:

req_id (RequestID)

Return type:

dict[str, int] | None

loading​

loading: dict[RequestID, LittleKVCacheBlock]

source

The checkpoint a connector hit is still loading, per request.

Hashless until its copy lands, so without this it reads as the live block.

page_size​

page_size: int = 0

source

the granularity a state can be committed at.

Type:

Tokens per page

release()​

release(req_id, replica_idx)

source

Frees every block the request holds, its successor too.

Parameters:

Return type:

None

resume()​

resume(ctx, replica_idx)

source

Returns the block this forward resumes from and the one it fills.

Filled from a checkpoint when the row holds one: what a prefix hit claimed or loaded, or the predecessor a checkpoint published. Filled with zeros when the request has processed nothing and matched no hit, since a drawn block holds whatever its last request wrote.

Empty once the request runs in a block it has written itself.

Parameters:

Return type:

Mapping[str, tuple[int | None, int]]

shrink_to_fit()​

shrink_to_fit(req_id, num_committed_blocks, replica_idx)

source

Refits the row to the committed blocks, keeping the live block last.

Frees every other block in the row. A loading checkpoint’s copy holds its own pin.

Parameters:

Return type:

None

successors​

successors: dict[RequestID, LittleKVCacheBlock]

source

The block each request’s next checkpoint continues the state in.

Held from admission to release while prefix caching is on, which is the second block blocks_to_reserve budgets. Drawing it in step instead would be a draw no admission check priced.