IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Python function

max_seq_len_fitting_in_cache

max_seq_len_fitting_in_cache()​

max.pipelines.kv_cache.max_seq_len_fitting_in_cache(params, available_cache_memory, is_di_enabled, model_name)

source

Returns the longest request the manager for this cache holds.

The cost of a request depends on the manager: the legacy pool charges every leaf a page per slot, Jenga charges each leaf only what it retains.

Returns:

The longest sequence, or None when no length exhausts the cache.

Parameters:

Return type:

int | None