For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python function
max_seq_len_fitting_in_cache
max_seq_len_fitting_in_cache()
max.pipelines.kv_cache.max_seq_len_fitting_in_cache(params, available_cache_memory, is_di_enabled, model_name)
Returns the longest request the manager for this cache holds.
The cost of a request depends on the manager: the legacy pool charges every leaf a page per slot, Jenga charges each leaf only what it retains.
-
Returns:
-
The longest sequence, or
Nonewhen no length exhausts the cache. -
Parameters:
-
- params (KVCacheParamInterface)
- available_cache_memory (int)
- is_di_enabled (bool)
- model_name (str)
-
Return type:
-
int | None