For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
matmul_dynamic_block_scaled_amd
def matmul_dynamic_block_scaled_amd[out_dtype: DType](c: TileTensor[out_dtype, Engine=c.Engine, address_space=c.address_space, linear_idx_type=c.linear_idx_type], a: TileTensor[.uint8, Engine=a.Engine, address_space=a.address_space, linear_idx_type=a.linear_idx_type], b: TileTensor[.uint8, Engine=b.Engine, address_space=b.address_space, linear_idx_type=b.linear_idx_type], a_scales: TileTensor[.float8_e8m0fnu, Engine=a_scales.Engine, address_space=a_scales.address_space, linear_idx_type=a_scales.linear_idx_type], b_scales: TileTensor[.float8_e8m0fnu, Engine=b_scales.Engine, address_space=b_scales.address_space, linear_idx_type=b_scales.linear_idx_type], ctx: DeviceContext)
Launches the AMD CDNA4 MXFP4 block-scaled matmul kernel.
Selects a BLOCK_N of 16 when the N dimension is divisible by 16,
otherwise falls back to 1, and enqueues
matmul_dynamic_block_scaled_amd_kernel over a 2D grid.
Parameters:
- out_dtype (
DType): Element type of the output matrixc.