IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo module

utils

Provides shared CPU matmul utilities including tile-consumer traits, kernel shape selection, and partial SIMD load/store helpers.

comptime values​

elementwise_compute_lambda_type​

comptime elementwise_compute_lambda_type = def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> SIMD[dtype, width]

elementwise_epilogue_type​

comptime elementwise_epilogue_type = def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None

ElementwiseComputeFn​

comptime ElementwiseComputeFn = def[dtype: DType, width: SIMDLength, *, alignment: Int](IndexList[Int(2)], SIMD[dtype, width]) -> SIMD[dtype, width] & RegisterPassable & ImplicitlyCopyable

Value-taking counterpart of elementwise_compute_lambda_type.

A unified closure passed as a runtime value carries the origins of its captures, so the buffers it reads stay alive until the kernel that calls it has launched. alignment has no default; pass alignment=1 where the legacy type relied on its default.

ElementwiseEpilogueFn​

comptime ElementwiseEpilogueFn = def[dtype: DType, width: SIMDLength, *, alignment: Int](IndexList[Int(2)], SIMD[dtype, width]) -> None & RegisterPassable & ImplicitlyCopyable

Value-taking counterpart of elementwise_epilogue_type.

The closure stores the output itself, so a kernel that receives one writes nothing to its output argument. Pass it to a named kernel with host_arg=. alignment has no default; pass alignment=1 where the legacy type relied on its default.

ElementwiseOutputComputeFn​

comptime ElementwiseOutputComputeFn = def[dtype: DType, width: SIMDLength, *, alignment: Int](IndexList[Int(2)], SIMD[dtype, width], SIMD[dtype, width]) -> SIMD[dtype, width] & RegisterPassable & ImplicitlyCopyable

An ElementwiseComputeFn that also receives the output's prior value.

Called as fn(idx, val, c_val), where c_val is the output tensor's value at idx before the kernel writes it, cast to dtype. The kernel reads c_val from its own output argument, so the closure never captures the output and cannot alias it.

no_epilogue_fn​

comptime no_epilogue_fn = _as_epilogue_fn[inflated$def[_no_epilogue_body[_, _, _]]](.__init__())

Stands in for an ElementwiseEpilogueFn argument on code paths that store the output directly.

Structs​

  • ​GemmShape: Helper class to unpack gemm dimension and layout.
  • ​InnerKernelID: Identifies the inner matmul kernel variant selected for a target.
  • ​KernelConfig: Static configuration of the matmul inner kernel.
  • ​MicroKernelShape: Record describing the inner kernel shape.
  • ​NullTileConsumer: No-op TileConsumer. Used as the default when no fusion is requested, and as the placeholder type when a kernel's tile_consumer Optional is None.
  • ​NullTileOperation: No-op TileOperation sentinel: parallel to NullTileConsumer. Used as the default TileOperationType so kernels without a fused op compile without callers having to spell out a placeholder.
  • ​SubMatmulConfig: Static configuration of sub-matrices in parallel matmul.

Traits​

  • ​TileConsumer: Trait for an epilogue operation which consumes a tile of data.
  • ​TileOperation: Non-terminal counterpart to TileConsumer: mutates a tile in place between MMA and store. Composes additively with capability subtraits like AuxLoading for ops that need pre-loaded data.

Functions​

Was this page helpful?