For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
generic_kv_cache_radd_dispatch
def generic_kv_cache_radd_dispatch[dtype: DType, collection_t: KVCollectionT, //, target: StringSpan[ImmStaticOrigin]](a: TileTensor[dtype, Engine=a.Engine, linear_idx_type=a.linear_idx_type], cache: collection_t, input_row_offsets: TileTensor[.uint32, Engine=input_row_offsets.Engine, linear_idx_type=input_row_offsets.linear_idx_type], batch_offset: UInt32, layer_idx: UInt32, ctx: DeviceContext)
Adds an input tensor elementwise into the paged KV cache in-place.
Splits the input tensor's last dimension into key and value halves and accumulates each half into the corresponding K or V cache slot, applying the batch offset and per-batch cache lengths to locate the target rows.
Parameters:
- dtype (
DType): Element type of the input tensoraand the KV cache entries (inferred). - collection_t (
KVCollectionT): ConcreteKVCollectionTtype of thecacheargument, used to recover the cache's static parameters and cache type (inferred). - target (
StringSpan[ImmStaticOrigin]): Compilation target string used to dispatch GPU versus CPU paths.
Args:
- a (
TileTensor[dtype, Engine=a.Engine, linear_idx_type=a.linear_idx_type]): Input tensor with shape (sum(seq_lens), 2 * hidden_size) where the first hidden_size columns target K and the rest target V. - cache (
collection_t): The collection storing the KVCache entries for this layer, retrieved via layer_idx. - input_row_offsets (
TileTensor[.uint32, Engine=input_row_offsets.Engine, linear_idx_type=input_row_offsets.linear_idx_type]): Tensor with shape (batch_size + 1,) denoting the start of each sequence along the ragged sequence dimension. - batch_offset (
UInt32): Offset added to the computed batch index to support batch slicing. - layer_idx (
UInt32): The index of the layer being executed, used to retrieve the KVCache objects from cache. - ctx (
DeviceContext): The call context pointer, passed by the graph compiler.