For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
o_smem_chunk_offset
def o_smem_chunk_offset[dtype: DType, swizzle_mode: TensorMapSwizzle, rows: Int](row: Int, chunk: Int) -> Int
Element offset of 16 B chunk chunk of output row row in the O staging tile a per-block swizzle_mode O-TMA store reads.
The tile is block-major [block, rows, K], K = swizzle_mode.bytes() // size_of[dtype], with swizzle_mode's XOR applied within each block, so
the 8 rows of one store phase land on 8 distinct 16 B bank groups for every
mode. chunk counts 16 B units along the row (8 bf16 or 4 f32 columns).
SWIZZLE_NONE reduces to one chunk per block, chunk * rows * K + row * K.
Returns: