For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo module
smallm_streaming_matmul
Small-M weight-streaming matmul over a preshuffled B for MI355X.
Serves decode-band vocab-class projections (huge N, K in the thousands, M <= 128 — lm_head shards and MTP eh_proj): the GEMM is >99% weight reads, so the kernel is built to stream B exactly once at HBM rate and spend nothing else. Each 16-column tile of C is owned by one block; its warps split K and run mfma 16x16x32 bf16 with the B fragment loaded straight from global memory.
B must be preshuffled into the fragment-major layout produced by
smallm_preshuffle_b (run once at weight-load time): a plain 16B load per
lane then lands a full warp instruction on 1KB of contiguous memory. Reading a
row-major B here is silently wrong — the layout is private to this pair, so
the op must be invoked explicitly, never via generic matmul dispatch.
A stays row-major [m, k] (contiguous, row stride == K) and is small enough
to be L2-resident. C is [m, n].
Functions
-
smallm_preshuffle_b: Launches the one-shot weight preshuffle (run at weight-load time). -
smallm_streaming_matmul: Runs the streaming matmul over asmallm_preshuffle_b-layout weight.