For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
grouped_matmul_amd_kernel_launcher
def grouped_matmul_amd_kernel_launcher[c_type: DType, a_type: DType, b_type: DType, LayoutC: TensorLayout, LayoutA: TensorLayout, LayoutB: TensorLayout, AOffsetsLayout: TensorLayout, ExpertIdsLayout: TensorLayout, transpose_b: Bool, config: MatmulConfig[a_type, b_type, c_type, transpose_b], c_engine: TensorEngine, a_engine: TensorEngine, b_engine: TensorEngine, a_offsets_engine: TensorEngine, expert_ids_engine: TensorEngine, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None](c_tensor: TileTensor[c_type, LayoutC, MutAnyOrigin, Engine=c_engine], a_tensor: TileTensor[a_type, LayoutA, MutAnyOrigin, Engine=a_engine], b_tensor: TileTensor[b_type, LayoutB, MutAnyOrigin, Engine=b_engine], a_offsets: TileTensor[.uint32, AOffsetsLayout, ImmUnsafeAnyOrigin, Engine=a_offsets_engine], expert_ids: TileTensor[.int32, ExpertIdsLayout, ImmUnsafeAnyOrigin, Engine=expert_ids_engine], num_active_experts: Int32)
Computes the AMD GPU grouped matmul by dispatching per-expert tiles through AMDMatmul, with separate zero-fill handling for inactive (expert_id == -1) blocks.
For active experts, delegates the per-tile matmul (and optional
elementwise epilogue) to AMDMatmul. For inactive experts, zeroes the
output row range and invokes the epilogue with zero values so that
LoRA-style inactive blocks still produce a defined output.