For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
naive_grouped_matmul_kernel
def naive_grouped_matmul_kernel[c_type: DType, a_type: DType, b_type: DType, CLayout: TensorLayout, ALayout: TensorLayout, BLayout: TensorLayout, AOffsetsLayout: TensorLayout, ExpertIdsLayout: TensorLayout, c_engine: TensorEngine, a_engine: TensorEngine, b_engine: TensorEngine, a_offsets_engine: TensorEngine, expert_ids_engine: TensorEngine, *, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None, a_plane_splits: IndexList[Int(2)] = Index[Int, Int](Int(0), Int(0))](c: TileTensor[c_type, CLayout, MutAnyOrigin, Engine=c_engine], a: TileTensor[a_type, ALayout, ImmUnsafeAnyOrigin, Engine=a_engine], b: TileTensor[b_type, BLayout, ImmUnsafeAnyOrigin, Engine=b_engine], a_offsets: TileTensor[.uint32, AOffsetsLayout, ImmUnsafeAnyOrigin, Engine=a_offsets_engine], expert_ids: TileTensor[.int32, ExpertIdsLayout, ImmUnsafeAnyOrigin, Engine=expert_ids_engine])
Computes one element per thread of the grouped matmul product C[a_offsets[z]:a_offsets[z+1], :] = A[...] @ B[expert_ids[z], :, :].T for each active expert z, with an optional elementwise epilogue.
Skips the matmul for expert == -1 (inactive LoRA blocks) but still
invokes the elementwise lambda when provided.