For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
with_mask_flash_attention_split_kv_cpu_shape
def with_mask_flash_attention_split_kv_cpu_shape(q: T, k: T, v: T, k_cache: T, v_cache: T, mask: T, scale: Float32) -> IndexList[T.rank]
Computes the output shape for the with_mask_flash_attention_split_kv_cpu graph op.
Returns:
IndexList[T.rank]