IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo module

smallm_streaming_matmul

Small-M weight-streaming matmul over a preshuffled B for MI355X.

Serves decode-band vocab-class projections (huge N, K in the thousands, M <= 128 — lm_head shards and MTP eh_proj): the GEMM is >99% weight reads, so the kernel is built to stream B exactly once at HBM rate and spend nothing else. Each 16-column tile of C is owned by one block; its warps split K and run mfma 16x16x32 bf16 with the B fragment loaded straight from global memory.

B must be preshuffled into the fragment-major layout produced by smallm_preshuffle_b (run once at weight-load time): a plain 16B load per lane then lands a full warp instruction on 1KB of contiguous memory. Reading a row-major B here is silently wrong — the layout is private to this pair, so the op must be invoked explicitly, never via generic matmul dispatch.

A stays row-major [m, k] (contiguous, row stride == K) and is small enough to be L2-resident. C is [m, n].

Functions​

Was this page helpful?