For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
grouped_matmul_mma_kernel
def grouped_matmul_mma_kernel[c_type: DType, a_type: DType, b_type: DType, N: Int, K: Int, linear_idx_type: DType, c_layout: TensorLayout, a_layout: TensorLayout, b_layout: TensorLayout, ao_layout: TensorLayout, ei_layout: TensorLayout, c_engine: TensorEngine, a_engine: TensorEngine, b_engine: TensorEngine, ao_engine: TensorEngine, ei_engine: TensorEngine](c: TileTensor[c_type, c_layout, MutAnyOrigin, Engine=c_engine], a: TileTensor[a_type, a_layout, ImmutAnyOrigin, Engine=a_engine], b: TileTensor[b_type, b_layout, ImmutAnyOrigin, Engine=b_engine], a_offsets: TileTensor[.uint32, ao_layout, ImmUnsafeAnyOrigin, Engine=ao_engine], expert_ids: TileTensor[.int32, ei_layout, ImmUnsafeAnyOrigin, Engine=ei_engine], first_group: UInt32, log2_grid_m: UInt32, log2_grid_n: UInt32)
Grouped GEMM: the dense AppleM5MatMul body on group first_group + block_idx.z.
Grid ((1 << log2_grid_m) * (1 << log2_grid_n), 1, groups), with the M
extent sized for the largest group; tiles past a smaller group's rows
early-return in the body. This is AppleM5MatMul.run with per-group views
in place of the whole-matrix operands.