IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

o_smem_chunk_offset

def o_smem_chunk_offset[dtype: DType, swizzle_mode: TensorMapSwizzle, rows: Int](row: Int, chunk: Int) -> Int

Element offset of 16 B chunk chunk of output row row in the O staging tile a per-block swizzle_mode O-TMA store reads.

The tile is block-major [block, rows, K], K = swizzle_mode.bytes() // size_of[dtype], with swizzle_mode's XOR applied within each block, so the 8 rows of one store phase land on 8 distinct 16 B bank groups for every mode. chunk counts 16 B units along the row (8 bf16 or 4 f32 columns). SWIZZLE_NONE reduces to one chunk per block, chunk * rows * K + row * K.

Returns:

Int

Was this page helpful?