IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo trait

InnerMatmulKernel

Trait for CPU matmul microkernels operating on pre-packed tiles.

Conforming types implement __inner_matmul__, which accumulates a (kernel_rows × TileN × TileK) block of the output matrix using a packed B tile in cache-friendly layout.

Implemented traits​

AnyType, Copyable, Deinitable, ImplicitlyCopyable, Movable

Provided methods​

__inner_matmul__​

def __inner_matmul__[kernel_rows: Int, kernel_cols: Int, simd_size: Int](self, c: TileTensor[Engine=c.Engine, address_space=c.address_space, linear_idx_type=c.linear_idx_type], a: TileTensor[Engine=a.Engine, address_space=a.address_space, linear_idx_type=a.linear_idx_type], b_packed: TileTensor[Engine=b_packed.Engine, address_space=b_packed.address_space, linear_idx_type=b_packed.linear_idx_type], global_offset: GemmShape, global_bound: GemmShape, tile_n_k: IndexList[Int(2)], skip_boundary_check: Bool)

Accumulates one packed B tile into the corresponding C tile.

Parameters:

  • ​kernel_rows (Int): Number of C rows the microkernel accumulates per tile.
  • ​kernel_cols (Int): Number of C columns the microkernel accumulates per tile.
  • ​simd_size (Int): SIMD vector width used by the inner accumulation.

Args:

Was this page helpful?