For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
sm100_heuristic_and_outliers_dispatch
def sm100_heuristic_and_outliers_dispatch[c_type: DType, a_type: DType, b_type: DType, ComputeFnType: ElementwiseComputeFn, //, transpose_b: Bool = True, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None, pdl_level: PDLLevel = PDLLevel(), has_epilogue_tensor: Bool = False, epilogue_is_1d: Bool = False, EpilogueEngine: TensorEngine = DefaultEngine, has_compute_fn: Bool = True](c: TileTensor[c_type, Engine=c.Engine, address_space=c.address_space, linear_idx_type=c.linear_idx_type], a: TileTensor[a_type, Engine=a.Engine, address_space=a.address_space, linear_idx_type=a.linear_idx_type], b: TileTensor[b_type, Engine=b.Engine, address_space=b.address_space, linear_idx_type=b.linear_idx_type], compute_fn: ComputeFnType, ctx: DeviceContext, epilogue_tensor: OptionalReg[TileTensor[c_type, Layout[TypeList[Int64, Int64](), TypeList[Int64, ComptimeInt[Int(1)]]()], ImmutAnyOrigin, Engine=EpilogueEngine]] = None) -> Int
Dispatches an SM100 matmul with a compute epilogue closure through the heuristic outlier config set.
Wraps select_and_launch_sm100_config with a launch callback that invokes
blackwell_matmul_tma_umma_warp_specialized directly, passing through the
elementwise epilogue lambda and compute_fn.
Parameters:
- c_type (
DType): Output element type (inferred). - a_type (
DType): Element type of the LHS operanda(inferred). - b_type (
DType): Element type of the RHS operandb(inferred). - ComputeFnType (
ElementwiseComputeFn): Type of the compute epilogue closure (inferred). - transpose_b (
Bool): Whetherbis stored transposed (defaults toTrue). - elementwise_lambda_fn (
Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None]): Optional epilogue applied to each output element (defaults toNone). - pdl_level (
PDLLevel): Programmatic dependent launch level for the dispatched kernel (defaults toPDLLevel()). - has_epilogue_tensor (
Bool): Whether an epilogue tensor is supplied for the TMA epilogue load path (defaults toFalse). - epilogue_is_1d (
Bool): Whether the epilogue tensor is treated as 1D rather than row-major 2D (defaults toFalse). - EpilogueEngine (
TensorEngine): Engine of the epilogue tensor (defaults toDefaultEngine[element_width=1]). - has_compute_fn (
Bool): Whether to applycompute_fn. When False,compute_fnis ignored (defaults toTrue).
Returns:
Int: DISPATCH_HIT when a kernel was launched, DISPATCH_MISS otherwise.