For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
naive_blockwise_scaled_fp8_matmul
def naive_blockwise_scaled_fp8_matmul[c_type: DType, a_type: DType, b_type: DType, a_scales_type: DType, b_scales_type: DType, //, *, BLOCK_DIM: Int = Int(16), transpose_b: Bool = False, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None, accum_type: DType = get_accum_type[c_type](), scales_granularity_mnk: Optional[IndexList[Int(3)]] = None](c: TileTensor[c_type, Engine=c.Engine, linear_idx_type=c.linear_idx_type], a: TileTensor[a_type, Engine=a.Engine, linear_idx_type=a.linear_idx_type], b: TileTensor[b_type, Engine=b.Engine, linear_idx_type=b.linear_idx_type], a_scales: TileTensor[a_scales_type, Engine=a_scales.Engine, linear_idx_type=a_scales.linear_idx_type], b_scales: TileTensor[b_scales_type, Engine=b_scales.Engine, linear_idx_type=b_scales.linear_idx_type], ctx: DeviceContext)
Dispatches the naive blockwise scaled FP8 matmul kernel on the GPU.
Enqueues naive_blockwise_scaled_fp8_matmul_kernel with a 2D grid of
BLOCK_DIM-sized tiles covering the M x N output.
Args:
- c (
TileTensor[c_type, Engine=c.Engine, linear_idx_type=c.linear_idx_type]): Rank-2 output accumulator tensor. - a (
TileTensor[a_type, Engine=a.Engine, linear_idx_type=a.linear_idx_type]): Rank-2 FP8 input matrix in K-major format. - b (
TileTensor[b_type, Engine=b.Engine, linear_idx_type=b.linear_idx_type]): Rank-2 FP8 weight matrix; K-major whentranspose_bis True, otherwise N-major. - a_scales (
TileTensor[a_scales_type, Engine=a_scales.Engine, linear_idx_type=a_scales.linear_idx_type]): Rank-2 per-block scales forain M-major format. - b_scales (
TileTensor[b_scales_type, Engine=b_scales.Engine, linear_idx_type=b_scales.linear_idx_type]): Rank-2 per-block scales forb; K-major whentranspose_bis True, otherwise N-major. - ctx (
DeviceContext): Device context used to enqueue the kernel.