IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

matmul_dispatch_sm100

def matmul_dispatch_sm100[c_type: DType, a_type: DType, b_type: DType, transpose_b: Bool = False, use_tf32: Bool = True, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None, elementwise_lambda_wrapper: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None, pdl_level: PDLLevel = PDLLevel()](c: TileTensor[c_type, Engine=c.Engine, address_space=c.address_space, linear_idx_type=c.linear_idx_type], a: TileTensor[a_type, Engine=a.Engine, address_space=a.address_space, linear_idx_type=a.linear_idx_type], b: TileTensor[b_type, Engine=b.Engine, address_space=b.address_space, linear_idx_type=b.linear_idx_type], ctx: DeviceContext)

Dispatches a 2D matmul to the appropriate SM100 (B200+) kernel.

Same as the compute_fn overload without a compute epilogue.

Parameters:

def matmul_dispatch_sm100[c_type: DType, a_type: DType, b_type: DType, ComputeFnType: ElementwiseComputeFn, //, transpose_b: Bool = False, use_tf32: Bool = True, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None, elementwise_lambda_wrapper: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None, pdl_level: PDLLevel = PDLLevel(), has_compute_fn: Bool = True](c: TileTensor[c_type, Engine=c.Engine, address_space=c.address_space, linear_idx_type=c.linear_idx_type], a: TileTensor[a_type, Engine=a.Engine, address_space=a.address_space, linear_idx_type=a.linear_idx_type], b: TileTensor[b_type, Engine=b.Engine, address_space=b.address_space, linear_idx_type=b.linear_idx_type], compute_fn: ComputeFnType, ctx: DeviceContext)

Dispatches a 2D matmul with a compute epilogue closure to the appropriate SM100 (B200+) kernel.

Routes the problem to GEMV for M=1 or N=1 shapes, to the IEEE-fp32 split-K GEMV for precise float32, or to the dtype-specific SM100 dispatcher (bf16, fp8, fp32) for general shapes, falling back to vendor BLAS when no Mojo SM100 config applies. In autotuning mode, launches a single compile-time-configured kernel from environment defines.

Only the SM100 GEMM paths apply compute_fn. The GEMV, small-MN, and vendor BLAS paths apply elementwise_lambda_wrapper instead, so when has_compute_fn is set the wrapper must apply the same epilogue and store the result.

Parameters:

Was this page helpful?