For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
naive_block_scaled_matmul_kernel
def naive_block_scaled_matmul_kernel[c_type: DType, a_type: DType, b_type: DType, a_scales_type: DType, b_scales_type: DType, accum_type: DType, a_layout: TensorLayout, b_layout: TensorLayout, c_layout: TensorLayout, a_scale_layout: TensorLayout, b_scale_layout: TensorLayout, scaling_kind: UMMAKind, SF_VECTOR_SIZE: Int, transpose_b: Bool = True, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None](c: TileTensor[c_type, c_layout, MutAnyOrigin], a: TileTensor[a_type, a_layout, ImmutAnyOrigin], b: TileTensor[b_type, b_layout, ImmutAnyOrigin], a_scales: TileTensor[a_scales_type, a_scale_layout, ImmutAnyOrigin], b_scales: TileTensor[b_scales_type, b_scale_layout, ImmutAnyOrigin], alpha: Float32)
Naive GPU kernel that emulates a block-scaled matmul using TCGEN scale factors.
Both A and B must be in K-major format with 5D TCGEN scale-factor layouts. Each thread accumulates one output element by iterating over K, applying per-block scale factors, and optionally invoking an elementwise epilogue lambda.
Parameters:
- c_type (
DType): Element type of the output matrixc. - a_type (
DType): Element type of the LHS input matrixa. - b_type (
DType): Element type of the RHS input matrixb. - a_scales_type (
DType): Element type of thea_scalesblock scale-factor tensor. - b_scales_type (
DType): Element type of theb_scalesblock scale-factor tensor. - accum_type (
DType): Element type used for the dot-product accumulator. - a_layout (
TensorLayout): Memory layout of the LHS input matrixa. - b_layout (
TensorLayout): Memory layout of the RHS input matrixb. - c_layout (
TensorLayout): Memory layout of the output matrixc. - a_scale_layout (
TensorLayout): Memory layout of thea_scalesblock scale-factor tensor. - b_scale_layout (
TensorLayout): Memory layout of theb_scalesblock scale-factor tensor. - scaling_kind (
UMMAKind):UMMAKindvariant selecting the block-scaled MMA instruction. - SF_VECTOR_SIZE (
Int): Number of elements covered by each block scale factor. - transpose_b (
Bool): Whetherbis stored transposed (defaults toTrue). - elementwise_lambda_fn (
Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None]): Optional epilogue lambda applied to the matmul result (defaults toNone).